vllm.model_executor.layers.rocm_paged_mxfp4_indexer
¶
Sparse attention indexer layers on aiter's paged MXFP4 kernels, for ROCm.
With indexer_kv_dtype="mxfp4" on gfx950, the model writes the indexer K
cache in the order aiter's MQA-logits kernel reads (see
rocm_paged_mxfp4_cache_layout), and the kernel scores the cache in place.
Requires RocmMxfp4IndexerMetadataBuilder metadata.
Classes:
-
RocmSparseAttnIndexer–SparseAttnIndexerwhose HIP path walks the paged MXFP4 cache with -
RocmSparseMQAIndexer–DeepSeek-V4.1's candidate-consuming indexer on aiter's paged MXFP4
RocmSparseAttnIndexer
¶
Bases: SparseAttnIndexer
SparseAttnIndexer whose HIP path walks the paged MXFP4 cache with
aiter's kernel. A two-level indexer's source layer also publishes its
candidate pool here.
Source code in vllm/model_executor/layers/rocm_paged_mxfp4_indexer.py
RocmSparseMQAIndexer
¶
Bases: Module
DeepSeek-V4.1's candidate-consuming indexer on aiter's paged MXFP4 MQA-logits kernel: the consumers score only the candidate pool the source layer published, read straight from the paged cache.
Built and called as SparseMQAIndexer is, so the model can take either.
Attributes:
-
weights_dtype–The scorer's weights and logits stay fp32: gfx950 has no packed bf16
Source code in vllm/model_executor/layers/rocm_paged_mxfp4_indexer.py
weights_dtype = torch.float32
class-attribute
instance-attribute
¶
The scorer's weights and logits stay fp32: gfx950 has no packed bf16 arithmetic, so its head reduce costs the same in either dtype.