Skip to content

vllm_omni.platforms.xpu

Modules:

Name Description
patch

XPU-specific patches for upstream vLLM behavior.

platform
profiler
utils
worker

XPUOmniPlatform

Bases: OmniPlatform, XPUPlatform

XPU/Intel GPU implementation of OmniPlatform.

Inherits all XPU-specific implementations from vLLM's XPUPlatform, and adds Omni-specific interfaces from OmniPlatform.

get_default_ir_op_priority classmethod

get_default_ir_op_priority(
    vllm_config: VllmConfig,
) -> IrOpPriorityConfig

Copied from upstream XPUPlatform with inductor-aware logic.

When inductor is active (compiling) use native as the default; otherwise prefer vllm_c where available.

get_default_stage_config_path classmethod

get_default_stage_config_path() -> str

get_device_count classmethod

get_device_count() -> int

get_device_memory classmethod

get_device_memory(
    device: device | None = None,
) -> tuple[int, int]

get_device_version classmethod

get_device_version() -> str | None

get_diffusion_attn_backend_cls classmethod

get_diffusion_attn_backend_cls(
    selected_backend: str | None,
    head_size: int,
    allow_trtllm_default: bool = False,
) -> str

get_free_memory classmethod

get_free_memory(device: device | None = None) -> int

get_omni_ar_worker_cls classmethod

get_omni_ar_worker_cls() -> str

get_omni_generation_worker_cls classmethod

get_omni_generation_worker_cls() -> str

get_profiler_cls classmethod

get_profiler_cls() -> str

Return XPU-specific profiler that handles XPU events.

get_torch_device classmethod

get_torch_device(local_rank: int | None = None) -> device

memory_reserved classmethod

memory_reserved(device: device | int | None = None) -> int

record_device_event classmethod

record_device_event() -> Event | None

Record an XPU event on the current stream to mark tensor readiness.

Deliberately a device-agnostic torch.Event rather than a torch.xpu.Event. The consumer (the async diffusion output thread) waits with torch.Stream.wait_event on a generic torch.Stream, and that C-level binding silently no-ops for a torch.xpu.Event instead of enqueuing the dependency — the side stream then starts its D2H copy while the compute stream is still writing the tensor, so the host reads a partially-written image (garbage rows at the bottom of the output). torch.Event dispatches through the accelerator hooks and the wait is honored, which is the actual fix.

supports_torch_inductor classmethod

supports_torch_inductor() -> bool

synchronize classmethod

synchronize() -> None