vllm.v1.attention.ops.rocm_paged_mxfp4_indexer
¶
Sparse attention indexer on aiter's paged MXFP4 MQA-logits kernel.
gfx950 only. The kernel reads the preshuffled paged indexer K cache in place,
so no layer gathers K into a contiguous buffer, prefill included. Any model
whose indexer supplies the inputs rocm_mxfp4_sparse_attn_indexer documents
can run the dense path. DeepSeek-V4.1's two-level indexer adds to it: the
candidate source takes its block maxima from the same walk that writes its
logits, and the candidate consumers can walk the candidate pool instead of
the whole context: the pool is resolved once per step and shared by all of
them.
Classes:
-
RocmPagedMxfp4CacheLayout–The byte order of an indexer K page, as the kernel reads it, and the
Functions:
-
build_rocm_mxfp4_decode_schedule–Work descriptors that even out a decode step, or None where the static
-
reserve_rocm_mxfp4_indexer_workspace–Profiling run: claim the decode logits workspace and the peak prefill
-
rocm_mxfp4_consumer_rows–Query rows per candidate-consumer launch. Its logits are [rows, pool]
-
rocm_mxfp4_decode_schedule_words–int32 words of the largest schedule a decode step can build, flattened or
-
rocm_mxfp4_sparse_attn_indexer–Dense indexer: every layer that scores the whole context, the
-
rocm_mxfp4_sparse_mqa_indexer–Candidate consumer: score only the source's pool where the builder's
-
rocm_paged_mxfp4_cache_layout–The K page layout from aiter's
cache_format, the only place vLLM
RocmPagedMxfp4CacheLayout
¶
Bases: NamedTuple
The byte order of an indexer K page, as the kernel reads it, and the shuffle pattern aiter's K cache op writes it with.
A page is cut into runs of n_per_tile tokens. Inside a run, each token's
packed values are split into d_per_tile-byte chunks along the head
dimension, and the run stores chunk 0 of every token, then chunk 1, and so
on: [chunk, token, byte]. Its e8m0 scales are stored as
[scale % scale_lanes, token, scale // scale_lanes].
Source code in vllm/v1/attention/ops/rocm_paged_mxfp4_indexer.py
_Layer
dataclass
¶
One indexer layer's inputs and outputs for this step.
Methods:
-
decode_rows–The first n token rows, one sequence each.
-
rows–Q, its scales and the weights of token rows [lo, hi), as
seqs
Source code in vllm/v1/attention/ops/rocm_paged_mxfp4_indexer.py
decode_rows(n)
¶
The first n token rows, one sequence each.
rows(lo, hi, seqs=1)
¶
Q, its scales and the weights of token rows [lo, hi), as seqs
sequences of next_n rows each.
Source code in vllm/v1/attention/ops/rocm_paged_mxfp4_indexer.py
_aiter()
cached
¶
The aiter module with the paged MXFP4 MQA-logits launcher, its schedule and cache_format.
_aiter_topk()
cached
¶
_kv_view(kv_cache, head_dim)
¶
The indexer cache as [pages, entries, 1, bytes]. The page stride is the block-major pool's, not the page's own size.
Source code in vllm/v1/attention/ops/rocm_paged_mxfp4_indexer.py
_remap_compact_topk(indices, positions, block)
¶
Candidate slots to context positions, in place: slot j sits in pool block j // block, which starts at positions[j // block].
Source code in vllm/v1/attention/ops/rocm_paged_mxfp4_indexer.py
_topk(logits, lengths, out, k)
¶
Top-k over each row's [0, lengths[row]); -1 pads rows shorter than k.
aiter's kernel is faster on the candidate top-k and from 128 rows on; below that, a wide decode step is faster on vLLM's. Shape alone decides, so a FULL graph replays the same choice at every context length.
Source code in vllm/v1/attention/ops/rocm_paged_mxfp4_indexer.py
build_rocm_mxfp4_decode_schedule(row_lens, num_heads, head_dim, page_entries, out, logits_width, native=None)
¶
Work descriptors that even out a decode step, or None where the static grid already fills the machine. Depends only on the rows' lengths, the cache geometry and the logits width, so one serves every layer of a group. The width sizes the slices, capped the way the static grid caps them.
Source code in vllm/v1/attention/ops/rocm_paged_mxfp4_indexer.py
reserve_rocm_mxfp4_indexer_workspace(hidden_states, logits_width, candidate_block_size=0, gather_block_size=0, num_candidate_cols=0)
¶
Profiling run: claim the decode logits workspace and the peak prefill logits, block scores included when the layer writes them, candidate lists when it gathers.
Source code in vllm/v1/attention/ops/rocm_paged_mxfp4_indexer.py
rocm_mxfp4_consumer_rows(num_candidate_cols)
¶
Query rows per candidate-consumer launch. Its logits are [rows, pool] fp32 whatever the context, so the logits budget alone sizes it.
Source code in vllm/v1/attention/ops/rocm_paged_mxfp4_indexer.py
rocm_mxfp4_decode_schedule_words(num_heads, head_dim, page_entries, next_n=1)
¶
int32 words of the largest schedule a decode step can build, flattened or not. The slice cap can take it past target_wgs, up to SCHED_SLOT_CAP.
Source code in vllm/v1/attention/ops/rocm_paged_mxfp4_indexer.py
rocm_mxfp4_sparse_attn_indexer(hidden_states, k_cache_prefix, kv_cache, q_values, q_scale, weights, topk_tokens, head_dim, max_model_len, topk_indices_buffer, compress_ratio=1, candidate_blocks=None, candidate_block_size=0, candidate_write=False)
¶
Dense indexer: every layer that scores the whole context, the
candidate source included. With candidate_blocks and not
candidate_write the scores are masked to the pool first, as the
shared path does.
The model supplies q_values [T, H, D // 2] packed e2m1, q_scale
holding each head's D // 32 ue8m0 bytes (one int32 per head at D = 128),
fp32 weights with the softmax and head scales folded in, and
kv_cache pages in rocm_paged_mxfp4_cache_layout order for H heads.
Source code in vllm/v1/attention/ops/rocm_paged_mxfp4_indexer.py
rocm_mxfp4_sparse_mqa_indexer(hidden_states, k_cache_prefix, kv_cache, q_values, q_scale, weights, topk_tokens, head_dim, max_model_len, topk_indices_buffer, compress_ratio, candidate_blocks, candidate_block_size, num_candidate_cols)
¶
Candidate consumer: score only the source's pool where the builder's length gate says it pays, else the dense walk masked to the pool.
Source code in vllm/v1/attention/ops/rocm_paged_mxfp4_indexer.py
rocm_paged_mxfp4_cache_layout(num_heads, head_dim, page_entries)
cached
¶
The K page layout from aiter's cache_format, the only place vLLM
reads it, so a change there cannot reach the writer unchecked.
Raises:
-
ValueError–aiter rejects the page size, or describes an order this layout cannot express.