Skip to content

vllm_omni.utils.device_copy

Host-to-device copies that do not block the host on queued GPU work.

A copy from pageable host memory (tensor.to("cuda"), torch.tensor(..., device="cuda"), torch.as_tensor(..., device="cuda")) makes the CPU wait until every kernel already queued on the stream has finished. In a per-step path under async scheduling that removes the CPU/GPU overlap the batch queue exists for. Staging the data in pinned memory and copying with non_blocking keeps the host running ahead; the caching host allocator keeps the pinned source alive until the copy's stream event completes.

index_to_device

index_to_device(
    values: Sequence[int],
    device: device | str,
    dtype: dtype = long,
) -> Tensor

A host index list as a device tensor, without a host sync.

to_device_nonblocking

to_device_nonblocking(
    tensor: Tensor, device: device | str
) -> Tensor

tensor.to(device) without a host sync when tensor is on the CPU and device is CUDA.