Skip to content

vllm_omni.engine.duplex.vad

Silero server VAD for the engine-resident duplex session.

The detector keeps the split upstream's serving-side server_vad module used, because the split is what makes one model serve many sessions: a backend scores one 512-sample frame and is otherwise stateless, and the per-stream state -- the partial frame, the model state, the endpoint counters -- belongs to :class:SileroStreamingVAD, one per session.

Two backends. :class:SileroVADBackend runs the pinned Silero v6.2 ONNX graph on CPU and is shared process-wide, which is why its state is passed in and out rather than held on the session object. :class:TorchSileroBackend loads the same model through the silero-vad package for environments that have no local ONNX artifact; torch keeps its state inside the module, so that one is per session.

Endpointing follows Silero v6.2's streaming hysteresis: activation uses the configured threshold, while a turn can only end on frames below max(threshold - 0.15, 0.01). Once a silence candidate exists, louder frames keep the elapsed-silence clock running but cannot themselves close the turn.

SILERO_VAD_FILENAME module-attribute

SILERO_VAD_FILENAME = 'silero_vad.onnx'

SILERO_VAD_MIN_THRESHOLD module-attribute

SILERO_VAD_MIN_THRESHOLD = 0.15

SILERO_VAD_REPO_ID module-attribute

SILERO_VAD_REPO_ID = 'istupakov/silero-vad-onnx'

SILERO_VAD_REVISION module-attribute

SILERO_VAD_REVISION = (
    "8b14476858ef240c50b3884bb38cc67290c1cc70"
)

SILERO_VAD_SHA256 module-attribute

SILERO_VAD_SHA256 = "1a153a22f4509e292a94e67d6f9b85e8deb25b4988682b7e174c65279d8788e3"

logger module-attribute

logger = init_logger(__name__)

ServerVADUnavailableError

Bases: RuntimeError

SileroStreamingVAD

config instance-attribute

config = config

scratch_bytes property

scratch_bytes: int

Bytes this detector holds between chunks (the partial frame only).

speech_active property

speech_active: bool

process

process(samples: ndarray) -> StreamingVADResult

process_base64

process_base64(
    audio: object, *, fmt: object, sample_rate_hz: object
) -> StreamingVADResult

reset

reset() -> None

Drop stream state while preserving the session-relative audio clock.

The clock survives so speech_start_ms after a barge-in still refers to the session's timeline; _stream_start_ms then stops the prefix padding reaching back into audio that was discarded.

SileroVADBackend

Shared ONNX Runtime Silero v6.2 detector running on CPU.

context_samples class-attribute instance-attribute

context_samples = 64

frame_samples class-attribute instance-attribute

frame_samples = _FRAME_SAMPLES

model_path instance-attribute

model_path = Path(model_path)

model_state_shape class-attribute instance-attribute

model_state_shape = (2, 1, 128)

sample_rate_hz class-attribute instance-attribute

sample_rate_hz = _SAMPLE_RATE_HZ

infer

infer(
    frame: ndarray, state: object
) -> tuple[float, object]

new_state

new_state() -> _SileroVADState

SileroVADBackendProvider

Resolve, verify and load one detector backend per engine process.

ONNX is preferred and shared; the torch package is the fallback and is built per call because its state is internal.

model_path instance-attribute

model_path = model_path

get

The shared ONNX backend, or a per-session torch one when it cannot be built.

An explicitly configured server_vad_model_path never falls back: the operator named that artifact, so a missing file or a missing ONNX Runtime is their error, not something to paper over with a different model.

SileroVADConfig dataclass

min_speech_duration_ms class-attribute instance-attribute

min_speech_duration_ms: int = 96

prefix_padding_ms class-attribute instance-attribute

prefix_padding_ms: int = 300

silence_duration_ms class-attribute instance-attribute

silence_duration_ms: int = 500

threshold class-attribute instance-attribute

threshold: float = 0.5

SpeechDetectorBackend

Bases: Protocol

Scores one frame. Per-stream state is created here but owned by the caller.

frame_samples instance-attribute

frame_samples: int

infer

infer(
    frame: ndarray, state: object
) -> tuple[float, object]

new_state

new_state() -> object

StreamingVADResult dataclass

is_speech instance-attribute

is_speech: bool

speech_active instance-attribute

speech_active: bool

speech_end_ms class-attribute instance-attribute

speech_end_ms: int | None = None

speech_probability class-attribute instance-attribute

speech_probability: float = 0.0

speech_start_ms class-attribute instance-attribute

speech_start_ms: int | None = None

speech_started class-attribute instance-attribute

speech_started: bool = False

speech_stopped class-attribute instance-attribute

speech_stopped: bool = False

TorchSileroBackend

The same model through the silero-vad package, for hosts with no ONNX artifact.

Torch keeps the recurrent state inside the module, so unlike the ONNX backend this one cannot be shared: each session gets its own instance and new_state only resets it.

frame_samples class-attribute instance-attribute

frame_samples = _FRAME_SAMPLES

sample_rate_hz class-attribute instance-attribute

sample_rate_hz = _SAMPLE_RATE_HZ

infer

infer(
    frame: ndarray, state: object
) -> tuple[float, object]

new_state

new_state() -> object