Skip to content

vllm_omni.entrypoints.duplex.warmup

Startup warmup probe for duplex /v1/realtime.

Moved out of api_server.py under P0.2 of #5227. The omni_run_server_worker scheduling and /v1/realtime hold-clients preamble stay in api_server.py until P0.3.

Video-required models default to one short audio-plus-frame turn so the four pipeline stages compile real shapes before a client is accepted. That is separate from the per-stage JIT kernel registry, which can finish in ~0.01 s without compiling anything.

DUPLEX_WARMUP_AUDIO_WAIT_S module-attribute

DUPLEX_WARMUP_AUDIO_WAIT_S = 150

DUPLEX_WARMUP_CLIENT_WAIT_S module-attribute

DUPLEX_WARMUP_CLIENT_WAIT_S = 180

logger module-attribute

logger = init_logger(__name__)

lookup_duplex_plugin

lookup_duplex_plugin(
    engine_client: object,
) -> object | None

startup_warmup_kind

startup_warmup_kind(
    plugin: object | None, warmup_frames: int
) -> str | None

Which throwaway realtime session to run before admitting clients.

video_turn is one short audio chunk plus one image. It is the default for models that require video, including when warmup_frames is 0. silent_frames is the older audio-only path and stays opt-in. A negative warmup_frames disables both. None means do not hold clients.

warmup_silence_unit

warmup_silence_unit(
    plugin: object | None,
) -> dict[str, object]

The append the silent-frames warmup sends: the plugin's silence unit, or the 16 kHz default.