Skip to content

vllm_omni.worker_v2.output_snapshot

Owned, dtype-grouped snapshots for asynchronous runner outputs.

PackedOutputSnapshot

Bases: dict

A normal payload mapping with an internal batched-copy plan.

The runner's snapshot-slot event protects the device slabs until D2H is complete. Host slabs are allocated per output, so downstream consumers can retain their views without depending on reuse of the device ring slot.

copy_to_cpu

copy_to_cpu(
    copy_tensor: Callable[[Tensor], Tensor],
) -> dict[str, Any]

pack_output_snapshot

pack_output_snapshot(
    payload: dict[str, Any],
    slot: dict[tuple[Any, ...], Tensor],
    *,
    max_buckets: int,
) -> PackedOutputSnapshot | None

Copy tensor leaves once per dtype/device, retaining their exact values.

No CUDA synchronization is introduced here. The caller must select the producer stream and wait for the previous consumer before reusing a slot.