vllm_omni.platforms.xpu ¶
Modules:
| Name | Description |
|---|---|
patch | XPU-specific patches for upstream vLLM behavior. |
platform | |
profiler | |
utils | |
worker | |
XPUOmniPlatform ¶
Bases: OmniPlatform, XPUPlatform
XPU/Intel GPU implementation of OmniPlatform.
Inherits all XPU-specific implementations from vLLM's XPUPlatform, and adds Omni-specific interfaces from OmniPlatform.
get_default_ir_op_priority classmethod ¶
Copied from upstream XPUPlatform with inductor-aware logic.
When inductor is active (compiling) use native as the default; otherwise prefer vllm_c where available.
get_device_memory classmethod ¶
get_diffusion_attn_backend_cls classmethod ¶
get_diffusion_attn_backend_cls(
selected_backend: str | None,
head_size: int,
allow_trtllm_default: bool = False,
) -> str
get_profiler_cls classmethod ¶
get_profiler_cls() -> str
Return XPU-specific profiler that handles XPU events.
record_device_event classmethod ¶
Record an XPU event on the current stream to mark tensor readiness.
Deliberately a device-agnostic torch.Event rather than a torch.xpu.Event. The consumer (the async diffusion output thread) waits with torch.Stream.wait_event on a generic torch.Stream, and that C-level binding silently no-ops for a torch.xpu.Event instead of enqueuing the dependency — the side stream then starts its D2H copy while the compute stream is still writing the tensor, so the host reads a partially-written image (garbage rows at the bottom of the output). torch.Event dispatches through the accelerator hooks and the wait is honored, which is the actual fix.