vllm_omni.entrypoints.duplex.warmup ¶
Startup warmup probe for duplex /v1/realtime.
Moved out of api_server.py under P0.2 of #5227. The omni_run_server_worker scheduling and /v1/realtime hold-clients preamble stay in api_server.py until P0.3.
Video-required models default to one short audio-plus-frame turn so the four pipeline stages compile real shapes before a client is accepted. That is separate from the per-stage JIT kernel registry, which can finish in ~0.01 s without compiling anything.
startup_warmup_kind ¶
Which throwaway realtime session to run before admitting clients.
video_turn is one short audio chunk plus one image. It is the default for models that require video, including when warmup_frames is 0. silent_frames is the older audio-only path and stays opt-in. A negative warmup_frames disables both. None means do not hold clients.