vllm_omni.diffusion.diffusion_kv.paged_attention_adapter ¶
DiffusionKVRowResolver module-attribute ¶
DiffusionKVRowResolver = Callable[
[str, int | None, str | None],
DiffusionPagedAttentionRowBinding,
]
DiffusionPagedAttentionAdapter ¶
Translate diffusion rows into vLLM metadata for an Omni paged backend.
activate ¶
activate(
batch: PreparedDiffusionPagedAttentionBatch,
) -> Iterator[DiffusionPagedAttentionAdapter]
invalidate_prepared_batches ¶
Invalidate native buffer views after BlockTable state changes.
prepare_batch ¶
prepare_batch(
rows: Sequence[DiffusionPagedAttentionRow],
) -> PreparedDiffusionPagedAttentionBatch
prepare_layer_context ¶
prepare_layer_context(
layer_name: str,
query: Tensor,
key: Tensor,
value: Tensor,
*,
omni_attn_metadata: Any | None = None,
) -> DiffusionPagedAttentionContext
DiffusionPagedAttentionContext dataclass ¶
One layer's native inputs for an Omni paged-backend invocation.
output_scatter_indices class-attribute instance-attribute ¶
DiffusionPagedAttentionLayerAdapter ¶
Bases: AttentionLayerBase
Register a diffusion layer with vLLM's native cache machinery.
This object deliberately does not subclass vllm.Attention. The latter owns a second execution path and would bypass Omni's sequence parallel pre/post hooks. The wrapper only supplies the small AttentionLayerBase contract needed by init_attn_backend and keeps the platform-native attention implementation/cache view available to the diffusion adapter.
head_size_v instance-attribute ¶
DiffusionPagedAttentionMetadata dataclass ¶
Runner-owned row layouts for one request-level denoise loop.
DiffusionPagedAttentionRow dataclass ¶
One logical BlockTable row participating in a paged attention call.
DiffusionPagedAttentionRowBinding dataclass ¶
DiffusionPagedAttentionRuntime ¶
Activate Runner-prepared rows lazily as the denoise loop advances.
PreparedDiffusionPagedAttentionBatch dataclass ¶
Native metadata shared by all paged attention layers in one forward.