vllm_omni.engine.duplex.vad ¶
Silero server VAD for the engine-resident duplex session.
The detector keeps the split upstream's serving-side server_vad module used, because the split is what makes one model serve many sessions: a backend scores one 512-sample frame and is otherwise stateless, and the per-stream state -- the partial frame, the model state, the endpoint counters -- belongs to :class:SileroStreamingVAD, one per session.
Two backends. :class:SileroVADBackend runs the pinned Silero v6.2 ONNX graph on CPU and is shared process-wide, which is why its state is passed in and out rather than held on the session object. :class:TorchSileroBackend loads the same model through the silero-vad package for environments that have no local ONNX artifact; torch keeps its state inside the module, so that one is per session.
Endpointing follows Silero v6.2's streaming hysteresis: activation uses the configured threshold, while a turn can only end on frames below max(threshold - 0.15, 0.01). Once a silence candidate exists, louder frames keep the elapsed-silence clock running but cannot themselves close the turn.
SILERO_VAD_REVISION module-attribute ¶
SILERO_VAD_SHA256 module-attribute ¶
ServerVADUnavailableError ¶
Bases: RuntimeError
SileroStreamingVAD ¶
scratch_bytes property ¶
scratch_bytes: int
Bytes this detector holds between chunks (the partial frame only).
process_base64 ¶
process_base64(
audio: object, *, fmt: object, sample_rate_hz: object
) -> StreamingVADResult
reset ¶
Drop stream state while preserving the session-relative audio clock.
The clock survives so speech_start_ms after a barge-in still refers to the session's timeline; _stream_start_ms then stops the prefix padding reaching back into audio that was discarded.
SileroVADBackend ¶
Shared ONNX Runtime Silero v6.2 detector running on CPU.
SileroVADBackendProvider ¶
Resolve, verify and load one detector backend per engine process.
ONNX is preferred and shared; the torch package is the fallback and is built per call because its state is internal.
get ¶
get() -> SpeechDetectorBackend
The shared ONNX backend, or a per-session torch one when it cannot be built.
An explicitly configured server_vad_model_path never falls back: the operator named that artifact, so a missing file or a missing ONNX Runtime is their error, not something to paper over with a different model.
SileroVADConfig dataclass ¶
SpeechDetectorBackend ¶
StreamingVADResult dataclass ¶
TorchSileroBackend ¶
The same model through the silero-vad package, for hosts with no ONNX artifact.
Torch keeps the recurrent state inside the module, so unlike the ONNX backend this one cannot be shared: each session gets its own instance and new_state only resets it.