vllm_omni.worker_v2.output_snapshot ¶
Owned, dtype-grouped snapshots for asynchronous runner outputs.
PackedOutputSnapshot ¶
Bases: dict
A normal payload mapping with an internal batched-copy plan.
The runner's snapshot-slot event protects the device slabs until D2H is complete. Host slabs are allocated per output, so downstream consumers can retain their views without depending on reuse of the device ring slot.
pack_output_snapshot ¶
pack_output_snapshot(
payload: dict[str, Any],
slot: dict[tuple[Any, ...], Tensor],
*,
max_buckets: int,
) -> PackedOutputSnapshot | None
Copy tensor leaves once per dtype/device, retaining their exact values.
No CUDA synchronization is introduced here. The caller must select the producer stream and wait for the previous consumer before reusing a slot.