Skip to content

vllm.models.hy_v4.nvidia

Modules:

  • attention

    MLA attention and lightning indexer for HY V4 (NVIDIA).

  • flashmla_sparse

    Sink-capable FlashMLA sparse backend for HY V4 (NVIDIA).

  • hc

    iHC (independent Hyper-Connections) layers for HY V4 (NVIDIA).

  • model

    Inference-only HY V4 model compatible with HuggingFace weights (NVIDIA).

  • moe

    Dense FFN and MoE blocks for HY V4 (NVIDIA).

  • mtp

    Multi-token prediction (MTP) head for HY V4 (NVIDIA).