Skip to content

vllm_omni.diffusion.offloader.base

logger module-attribute

logger = init_logger(__name__)

OffloadBackend

Bases: ABC

Base class for CPU offload backends

config instance-attribute

config = config

device instance-attribute

device = device

enabled instance-attribute

enabled = False

disable abstractmethod

disable() -> None

Disable offloading and cleanup resources.

Removes all registered hooks. Does NOT move modules back to original devices (caller responsible for that).

enable abstractmethod

enable(pipeline: Module) -> None

Enable offloading on the pipeline.

Discovers modules, moves them to appropriate devices, and registers forward hooks for swapping/prefetching.

Parameters:

Name Type Description Default
pipeline Module

Diffusion pipeline model (e.g., Wan22Pipeline)

required

is_enabled

is_enabled() -> bool

shutdown

shutdown() -> None

Release offload resources at process exit.

Backends may skip work that only matters for a later enable.

OffloadConfig dataclass

components class-attribute instance-attribute

components: frozenset[str] | None = None

dlo_host_registration_limit_gib class-attribute instance-attribute

dlo_host_registration_limit_gib: float = 0.0

dlo_resident_layers class-attribute instance-attribute

dlo_resident_layers: int = 0

dlo_transfers class-attribute instance-attribute

dlo_transfers: dict[str, DLOTransfer] | None = None

dlo_use_allgather class-attribute instance-attribute

dlo_use_allgather: bool = True

dp_size class-attribute instance-attribute

dp_size: int = 1

pin_cpu_memory class-attribute instance-attribute

pin_cpu_memory: bool = True

strategy instance-attribute

strategy: OffloadStrategy

use_hsdp class-attribute instance-attribute

use_hsdp: bool = False

from_od_config classmethod

from_od_config(
    od_config: OmniDiffusionConfig,
) -> OffloadConfig

Extract and validate offload settings from OmniDiffusionConfig.

diffusion_offload_config is the canonical public selector. The historical enable_*_offload booleans remain compatibility aliases; ambiguous combinations fail instead of using silent precedence.

The dp_size is automatically derived from parallel_config — it is NOT a user-configurable parameter. The distributed layerwise offload works with whatever DP/SP parallelism is already set up.

Parameters:

Name Type Description Default
od_config OmniDiffusionConfig

OmniDiffusionConfig with offload settings

required

Returns:

Type Description
OffloadConfig

OffloadConfig with validated settings

offloads

offloads(component: str) -> bool

offloads_encoder

offloads_encoder(
    name: str, plan: OffloadPlan | None = None
) -> bool

Return whether the selector covers a discovered encoder path.

Plans declare non-standard encoder names explicitly. The name-based fallback preserves compatibility with pipelines that predate OffloadPlan.

should_offload_encoder

should_offload_encoder(
    name: str, plan: OffloadPlan | None = None
) -> bool

Apply explicit selection while preserving the legacy encoder topology.

transfer_for

transfer_for(component: str) -> DLOTransfer

uses_allgather

uses_allgather(component: str) -> bool

SupportsModelCpuOffload

Bases: Protocol

Pipeline-owned lifecycle for model-level CPU offload.

Pipelines with non-forward component entry points (for example VAE decode_latent methods) need to activate those stages explicitly, so generic forward-hook discovery cannot manage their full lifecycle.

disable_omni_model_cpu_offload

disable_omni_model_cpu_offload() -> None

enable_omni_model_cpu_offload

enable_omni_model_cpu_offload(
    *,
    device: device,
    pin_memory: bool,
    use_hsdp: bool,
    offload_components: frozenset[str] | None = None,
) -> None

run_cleanup_steps

run_cleanup_steps(
    steps: Iterable[tuple[str, Callable[[], None]]],
) -> BaseException | None

Run every cleanup step and return the first failure, if any.