vllm_omni.diffusion.sched ¶
Modules:
| Name | Description |
|---|---|
base_scheduler | |
interface | |
request_scheduler | |
sigma_schedule | |
step_scheduler | |
BaseScheduler ¶
Bases: ABC
Shared queue/state bookkeeping for diffusion schedulers.
finish_requests ¶
finish_requests(
request_ids: str | list[str],
status: DiffusionRequestStatus,
) -> None
get_admission_wait_decision ¶
Return the admission-delay policy for the next scheduling wave.
get_diffusion_kv_cleanup_targets ¶
initialize ¶
initialize(
od_config: OmniDiffusionConfig,
*,
kv_cache_config: KVCacheConfig | None = None,
scheduler_block_size: int | None = None,
hash_block_size: int | None = None,
kv_vllm_config: VllmConfig | None = None,
) -> None
native_kv_poll_output ¶
native_kv_poll_output(
*, drain_request_ids: list[str] | None = None
) -> DiffusionSchedulerOutput | None
Poll after compute; cancellation/close may wait for selected loads.
pending_finished_request_ids ¶
Finished requests whose state the engine has not consumed yet.
release_kv_drains ¶
Called only after all ranks completed and Worker row cleanup succeeded.
should_end_admission_wait ¶
should_end_admission_wait(
decision: _AdmissionWaitDecision,
*,
now: float,
stable_since: float,
) -> bool
Return whether an active admission delay should end.
update_from_output abstractmethod ¶
update_from_output(
sched_output: DiffusionSchedulerOutput,
output: BaseRunnerOutput,
) -> set[str]
update_kv_connector_output ¶
CachedRequestData dataclass ¶
Cached diffusion requests that only need their request ids resent.
DMD2SigmaSchedule dataclass ¶
Continuous rectified-flow positions pinned by a distilled checkpoint.
A DMD2 student only ever sees the few noise levels it was trained on, so a distilled release ships the exact positions instead of letting the server derive a uniform schedule from num_inference_steps.
This is deliberately distinct from vllm_omni.diffusion.models.dmd2.DMD2Config.denoising_timesteps, which carries integer scheduler timesteps for scheduler-backed pipelines. Here the entries are continuous positions in [0, 1] that still need a per-modality time shift applied, which is what lets one schedule drive several coupled modalities at different shift scales.
num_inference_steps property ¶
num_inference_steps: int
Denoising steps, i.e. one per interval between sigma boundaries.
from_metadata classmethod ¶
from_metadata(
metadata: Mapping[str, Any],
*,
key: str = BASE_SCHEDULE_KEY,
) -> DMD2SigmaSchedule | None
Read a schedule from checkpoint metadata.
An absent key means the release is not distilled and keeps the legacy uniform schedule. An explicitly empty value is a malformed contract and is rejected rather than silently falling back.
DiffusionSchedulerOutput dataclass ¶
Output of a single scheduling cycle.
kv_connector_metadata class-attribute instance-attribute ¶
kv_finished_request_ids class-attribute instance-attribute ¶
kv_prefetch_connector_metadata class-attribute instance-attribute ¶
kv_prefetch_request_ids class-attribute instance-attribute ¶
kv_required_request_ids class-attribute instance-attribute ¶
kv_transfer_request_ids class-attribute instance-attribute ¶
scheduled_request_ids cached property ¶
All scheduled request ids in this cycle, including both new and cached ones.
KVPrefetchJob ¶
NewRequestData dataclass ¶
Payload for a newly scheduled diffusion request.
Carries the already-initialized request object so executors and workers do not re-run OmniDiffusionRequest.__post_init__ and mutate sentinel-based fields like guidance_scale_provided.
diffusion_kv_metadata class-attribute instance-attribute ¶
diffusion_kv_metadata: DiffusionKVMetadata | None = None
from_state classmethod ¶
from_state(
state: SchedulerRequestState,
*,
diffusion_kv_metadata: DiffusionKVMetadata
| None = None,
) -> NewRequestData
RequestScheduler ¶
Bases: BaseScheduler
Scheduler for static request waves, including admission coalescing.
get_admission_wait_decision ¶
should_end_admission_wait ¶
should_end_admission_wait(
decision: _AdmissionWaitDecision,
*,
now: float,
stable_since: float,
) -> bool
update_from_output ¶
update_from_output(
sched_output: DiffusionSchedulerOutput,
output: BaseRunnerOutput,
) -> set[str]
SchedulerInterface ¶
Bases: BaseScheduler
Deprecated compatibility base for custom scheduler injection.
Prefer subclassing :class:BaseScheduler directly. Subclassing this name still works but emits a :class:DeprecationWarning.
SchedulerRequestState dataclass ¶
Scheduler-owned state for one queued OmniDiffusionRequest.
diffusion_kv_requests class-attribute instance-attribute ¶
diffusion_kv_requests: tuple[DiffusionKVRequest, ...] = ()
sampling_params_key class-attribute instance-attribute ¶
sampling_params_key: (
StepBatchSamplingParamsKey
| RequestBatchSamplingParamsKey
| None
) = None
status class-attribute instance-attribute ¶
status: DiffusionRequestStatus = (
DiffusionRequestStatus.WAITING
)
StepBatchSamplingParamsKey dataclass ¶
Denoise step level Batch-compatibility key derived from OmniDiffusionSamplingParams.
Only requests with the same key can be batched together. Fields not included here are treated as request-local and do not participate in the current homogeneous batching policy.
StepScheduler ¶
Bases: BaseScheduler
Scheduler that advances each request by one denoise step per update.
update_from_output ¶
update_from_output(
sched_output: DiffusionSchedulerOutput,
output: BaseRunnerOutput,
) -> set[str]