vllm_omni.diffusion.offloader.base ¶
OffloadBackend ¶
Bases: ABC
Base class for CPU offload backends
disable abstractmethod ¶
Disable offloading and cleanup resources.
Removes all registered hooks. Does NOT move modules back to original devices (caller responsible for that).
enable abstractmethod ¶
Enable offloading on the pipeline.
Discovers modules, moves them to appropriate devices, and registers forward hooks for swapping/prefetching.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pipeline | Module | Diffusion pipeline model (e.g., Wan22Pipeline) | required |
shutdown ¶
Release offload resources at process exit.
Backends may skip work that only matters for a later enable.
OffloadConfig dataclass ¶
dlo_host_registration_limit_gib class-attribute instance-attribute ¶
dlo_host_registration_limit_gib: float = 0.0
dlo_transfers class-attribute instance-attribute ¶
dlo_transfers: dict[str, DLOTransfer] | None = None
from_od_config classmethod ¶
from_od_config(
od_config: OmniDiffusionConfig,
) -> OffloadConfig
Extract and validate offload settings from OmniDiffusionConfig.
diffusion_offload_config is the canonical public selector. The historical enable_*_offload booleans remain compatibility aliases; ambiguous combinations fail instead of using silent precedence.
The dp_size is automatically derived from parallel_config — it is NOT a user-configurable parameter. The distributed layerwise offload works with whatever DP/SP parallelism is already set up.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
od_config | OmniDiffusionConfig | OmniDiffusionConfig with offload settings | required |
Returns:
| Type | Description |
|---|---|
OffloadConfig | OffloadConfig with validated settings |
offloads_encoder ¶
offloads_encoder(
name: str, plan: OffloadPlan | None = None
) -> bool
Return whether the selector covers a discovered encoder path.
Plans declare non-standard encoder names explicitly. The name-based fallback preserves compatibility with pipelines that predate OffloadPlan.
should_offload_encoder ¶
should_offload_encoder(
name: str, plan: OffloadPlan | None = None
) -> bool
Apply explicit selection while preserving the legacy encoder topology.