vllm.v1.attention.backends.mla.rocm_paged_mxfp4_indexer
¶
Indexer metadata for the ROCm paged MXFP4 path.
Extends the dense indexer metadata with what aiter's paged MXFP4 MQA-logits
kernel needs to read the cache in place: each prefill chunk as its requests,
packed through query_start_loc into one launch, and the decode rows as
next_n-row sequences on uniform steps. RocmMxfp4IndexerMetadataBuilder plans
this for any model's indexer, and a two-level indexer's candidate blocks once
a subclass sets their size. DeepSeek-V4.1's builder does, and with
AttentionConfig.indexer_sparse_logits adds the per-step state the
candidate consumers share.
Classes:
-
DeepseekV41RocmMxfp4IndexerMetadataBuilder–DeepSeek-V4.1's two-level indexer: sets the candidate block size the
-
RocmMxfp4GatherLaunch–Query rows [token_start, token_end) of one request, which the candidate
-
RocmMxfp4IndexerMetadata– -
RocmMxfp4IndexerMetadataBuilder–Dense indexer metadata plus the in-place paged launches' layout.
-
RocmMxfp4NativeDecode–A decode step whose requests all have next_n query rows, as the
-
RocmMxfp4PrefillPlan–One prefill chunk as the requests the kernel launches on, all in one
Functions:
-
native_decode–The step as next_n-row sequences, or None when a request has fewer rows.
-
plan_gather_launches–The candidate consumers' launches over the chunks that gather.
-
plan_prefill_chunks–Split each chunk into its requests.
-
split_prefill_chunks–(request slice, query slice) chunks, as the base chunker returns them.
DeepseekV41RocmMxfp4IndexerMetadataBuilder
¶
Bases: RocmMxfp4IndexerMetadataBuilder
DeepSeek-V4.1's two-level indexer: sets the candidate block size the
base plans with and, with indexer_sparse_logits, adds the consumers'
gathers.
Source code in vllm/v1/attention/backends/mla/rocm_paged_mxfp4_indexer.py
RocmMxfp4GatherLaunch
dataclass
¶
Query rows [token_start, token_end) of one request, which the candidate consumers score against the pool in one launch.
Attributes:
-
block_table(Tensor) –[1, max_blocks] int32, the request's block table row.
-
context_len(Tensor) –[1] int32 compressed context of the request.
-
pool(tuple[dict, Tensor] | None) –The resolved candidate lists, built by the first consumer layer and
-
row_ends(Tensor) –[rows] int32 exclusive compressed key bound of each row.
Source code in vllm/v1/attention/backends/mla/rocm_paged_mxfp4_indexer.py
block_table
instance-attribute
¶
[1, max_blocks] int32, the request's block table row.
context_len
instance-attribute
¶
[1] int32 compressed context of the request.
pool = None
class-attribute
instance-attribute
¶
The resolved candidate lists, built by the first consumer layer and reused by the rest.
row_ends
instance-attribute
¶
[rows] int32 exclusive compressed key bound of each row.
RocmMxfp4IndexerMetadata
dataclass
¶
Bases: DeepseekV32IndexerMetadata
Attributes:
-
decode_block_ends(Tensor | None) –[rows] int32 candidate blocks each decode row sees, for the source's
-
decode_gather(tuple[dict, Tensor] | None) –The pool the first consumer resolved for the decode rows.
-
decode_native(RocmMxfp4NativeDecode | None) –The same rows as next_n-row sequences, for the dense launches.
-
decode_row_lens(Tensor | None) –[rows] int32 compressed context of each decode query row.
-
decode_schedule(Tensor | None) –The dense decode launches' work schedule, built once for the step.
-
decode_use_gather(bool) –Whether the candidate consumers gather for the decode rows.
-
gather_launches(list[RocmMxfp4GatherLaunch]) –The prefill rows of the chunks that gather, as the consumers launch
-
prefill_plans(list[RocmMxfp4PrefillPlan]) –Parallel to
prefill.chunks.
Source code in vllm/v1/attention/backends/mla/rocm_paged_mxfp4_indexer.py
decode_block_ends = None
class-attribute
instance-attribute
¶
[rows] int32 candidate blocks each decode row sees, for the source's pool.
decode_gather = None
class-attribute
instance-attribute
¶
The pool the first consumer resolved for the decode rows.
decode_native = None
class-attribute
instance-attribute
¶
The same rows as next_n-row sequences, for the dense launches.
decode_row_lens = None
class-attribute
instance-attribute
¶
[rows] int32 compressed context of each decode query row.
decode_schedule = None
class-attribute
instance-attribute
¶
The dense decode launches' work schedule, built once for the step.
decode_use_gather = False
class-attribute
instance-attribute
¶
Whether the candidate consumers gather for the decode rows.
gather_launches = field(default_factory=list)
class-attribute
instance-attribute
¶
The prefill rows of the chunks that gather, as the consumers launch them.
prefill_plans = field(default_factory=list)
class-attribute
instance-attribute
¶
Parallel to prefill.chunks.
RocmMxfp4IndexerMetadataBuilder
¶
Bases: DeepseekV32IndexerMetadataBuilder
Dense indexer metadata plus the in-place paged launches' layout.
Requirements are checked here, at engine start.
Source code in vllm/v1/attention/backends/mla/rocm_paged_mxfp4_indexer.py
304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 | |
RocmMxfp4NativeDecode
dataclass
¶
A decode step whose requests all have next_n query rows, as the kernel's sequences: a workgroup then walks a KV tile once for up to next_n rows instead of once per row.
Attributes:
-
block_table(Tensor) –[requests, max_blocks] int32, each request's first flattened row.
-
context_lens(Tensor) –[requests] int32 compressed context of each request.
Source code in vllm/v1/attention/backends/mla/rocm_paged_mxfp4_indexer.py
RocmMxfp4PrefillPlan
dataclass
¶
One prefill chunk as the requests the kernel launches on, all in one launch: the kernel finds each row's request through query_start_loc, and a request's rows share its KV tiles.
Attributes:
-
block_ends(Tensor | None) –[rows] int32 candidate blocks each row sees, for the source's pool.
-
context_lens(Tensor) –[chunk.num_reqs] int32 compressed context of each request.
-
first_request(int) –The step's index of the chunk's first request.
-
query_start_loc(Tensor) –[chunk.num_reqs + 1] int32 row offsets of the chunk's requests, local to
-
requests(list[tuple[int, int, int]]) –(first row, end row, request) per request with rows in the chunk; rows
-
row_ends(Tensor) –[rows] int32 exclusive compressed key bound of each row.
-
width(int) –Logits columns: an upper bound on the chunk's compressed contexts.
Source code in vllm/v1/attention/backends/mla/rocm_paged_mxfp4_indexer.py
block_ends = None
class-attribute
instance-attribute
¶
[rows] int32 candidate blocks each row sees, for the source's pool.
context_lens
instance-attribute
¶
[chunk.num_reqs] int32 compressed context of each request.
first_request
instance-attribute
¶
The step's index of the chunk's first request.
query_start_loc
instance-attribute
¶
[chunk.num_reqs + 1] int32 row offsets of the chunk's requests, local to the chunk.
requests
instance-attribute
¶
(first row, end row, request) per request with rows in the chunk; rows
count from chunk.token_start, requests index chunk.block_table.
row_ends
instance-attribute
¶
[rows] int32 exclusive compressed key bound of each row.
width
instance-attribute
¶
Logits columns: an upper bound on the chunk's compressed contexts.
native_decode(row_lens, row_block_table, query_lens, next_n, context_lens)
¶
The step as next_n-row sequences, or None when a request has fewer rows.
row_lens and row_block_table are the flattened rows' bounds and
block table, query_lens the decode requests' query lengths. Trailing
empty requests are cudagraph padding: they keep their next_n rows, with no
context. Only the step's shape decides, so a FULL graph captured on a
uniform batch replays the same launch on a padded one.
Source code in vllm/v1/attention/backends/mla/rocm_paged_mxfp4_indexer.py
plan_gather_launches(chunks, plans, max_rows)
¶
The candidate consumers' launches over the chunks that gather.
A consumer's logits are [rows, pool] however long the context, so the
dense split, sized for the context, would launch it far more often than
its memory needs. Each request's rows go back together across the chunks
that cut them, then out again at most max_rows a launch.
Source code in vllm/v1/attention/backends/mla/rocm_paged_mxfp4_indexer.py
plan_prefill_chunks(chunks, query_start_loc, seq_lens_cpu, context_lens, compress_ratio, candidate_block_size, min_gather_width, query_start_loc_device)
¶
Split each chunk into its requests.
query_start_loc and seq_lens_cpu are the step's host copies (the
latter an upper bound), context_lens the device compressed lengths.
A chunk covers whole requests, or a query slice of one request; either
way it launches once, packed through the chunk's part of the device
query_start_loc.
Source code in vllm/v1/attention/backends/mla/rocm_paged_mxfp4_indexer.py
split_prefill_chunks(context_lens, query_lens, max_logits_bytes, request_offset=0)
¶
(request slice, query slice) chunks, as the base chunker returns them.
Nothing gathers K here, so the requests' contexts do not add up: a chunk's
logits are [rows, widest compressed context] fp32, at most
max_logits_bytes. A request too wide to launch whole is cut on its
query rows.