vllm.models.hy_v4.nvidia ¶
Modules:
-
attention–MLA attention and lightning indexer for HY V4 (NVIDIA).
-
flashmla_sparse–Sink-capable FlashMLA sparse backend for HY V4 (NVIDIA).
-
hc–iHC (independent Hyper-Connections) layers for HY V4 (NVIDIA).
-
model–Inference-only HY V4 model compatible with HuggingFace weights (NVIDIA).
-
moe–Dense FFN and MoE blocks for HY V4 (NVIDIA).
-
mtp–Multi-token prediction (MTP) head for HY V4 (NVIDIA).