vllm_omni.protocol.realtime ¶
The OpenAI Realtime wire codec, shared by every vLLM-Omni Realtime surface.
What this is¶
The model-agnostic half of WS /v1/realtime: how a client event is parsed, which audio formats and session fields are accepted, what a conversation item means, how audio is decoded and resampled, and the shape of the error envelope (the codes inside it are a consumer's own --- ours are in vllm_omni.protocol.duplex.errors). It is a library of pure functions and value objects.
What it is not¶
It owns no session. It does not decide when a response starts, what a model is prompted with, whether a turn is over, or where committed audio goes. Those are runtime decisions, and after PR #7413 a duplex session's runtime decisions all belong to vllm_omni.engine.duplex --- the session runner, the typed DuplexCommand / DuplexEvent boundary and the DuplexModelPlugin seam. This package sits under that: the duplex engine binds the codec to its own command and event vocabulary in vllm_omni.engine.duplex.realtime_commands and vllm_omni.engine.duplex.realtime_events.
Why it is separate (RFC #6592, P0a)¶
So that a second Realtime surface --- a Qwen3-Omni GA profile, say --- can reuse the parsing, validation, format negotiation and error envelope without adopting the duplex session control plane (lease manager, commit policy, model channel), and without a second copy of the codec drifting away from this one. A consumer states what it can serve through :class:~vllm_omni.protocol.realtime.capabilities.RealtimeProtocolCapabilities and keeps its own session state.
Known duplex-shaped defaults¶
This package is model-agnostic in its dependencies, not yet in every default. Two constants still carry the assumptions of its first consumer and a second one will have to parameterize them:
- :func:
~vllm_omni.protocol.realtime.audio.convert_input_audio_with_rateresamples to 16 kHz, which is MiniCPM-o's feature-extractor rate. A GA client sending 24 kHz is converted down rather than served at its own rate. - :func:
~vllm_omni.protocol.realtime.items.validate_realtime_video_framesencodes the MiniCPM-o duplex camera contract (at most two frames per append,max_slice_numsnot selectable on the wire).
Both are behaviour this extraction deliberately did not change (RFC #6592 P0a is behaviour-preserving); making them per-consumer is P0b/P1 work.
Dependency rule¶
Nothing here may import vllm_omni.engine, vllm_omni.entrypoints, vllm_omni.model_executor, vllm_omni.worker or vllm_omni.clients; tests/protocol/realtime/test_protocol_import_boundary.py asserts it.
Modules:
| Name | Description |
|---|---|
audio | |
audio_input | Decoding one |
capabilities | What a Realtime consumer can serve, and the session check that reads it. |
commands | OpenAI Realtime client events as typed commands. |
errors | The OpenAI Realtime error envelope. |
events | OpenAI Realtime server events: typed classes and their wire rendering. |
formats | Realtime audio-format negotiation: what a client may declare and how it is read. |
items | Realtime conversation items: shape, truncation and the transcript inside one. |
session | Reading the Realtime |
REALTIME_INPUT_AUDIO_FORMATS module-attribute ¶
REALTIME_INPUT_AUDIO_FORMATS = {
"pcm16",
"pcm_s16le",
"s16le",
"pcm_f32le",
"g711_ulaw",
"g711_alaw",
}
REALTIME_INPUT_HINT_KEYS module-attribute ¶
REALTIME_INPUT_HINT_KEYS = (
"duration_ms",
"audio_duration_ms",
"audio_start_ms",
"audio_end_ms",
"is_speech",
"speech",
"speech_probability",
"vad",
"overlap_action",
"overlap",
"force_barge_in",
"force_listen",
"text",
"transcript",
)
REALTIME_OUTPUT_AUDIO_FORMATS module-attribute ¶
REALTIME_OUTPUT_AUDIO_FORMATS = {
"pcm16",
"pcm_s16le",
"s16le",
"wav",
"pcm",
"g711_ulaw",
"g711_alaw",
}
RealtimeAudioAppend dataclass ¶
One decoded input_audio_buffer.append.
audio is raw bytes in format at sample_rate_hz --- after :func:decode_audio_append that is 16 kHz pcm_f32le for every input format the codec converts.
RealtimeCommand dataclass ¶
One decoded OpenAI Realtime client event.
wire_type is the event type the client sent; the typed fields are what survived decoding, so a consumer reads them instead of re-parsing JSON.
There is deliberately no rendering method here. A server receives commands, it does not emit them, and how a runtime represents one internally is its own business --- the duplex engine renders its mailbox dictionary in vllm_omni.engine.duplex.commands.
RealtimeEvent dataclass ¶
Base of every public session event.
RealtimeInputDefaults dataclass ¶
Session-level wire defaults an append may omit (derived from the session object).
with_session_payload ¶
with_session_payload(
session_payload: Mapping[str, object],
) -> RealtimeInputDefaults
Return defaults updated from a Realtime session object (session.update).
RealtimeProtocolCapabilities dataclass ¶
One consumer's answer to what a session object may ask for.
The defaults are permissive, not restrictive: every format the codec can decode, and no turn-detection validation at all. A consumer that leaves validate_turn_detection at None therefore accepts whatever turn_detection the session object carries --- including values it cannot actually serve, such as semantic_vad. A consumer that supports only some turn-detection modes (or none) must supply its own validator; see vllm_omni.engine.duplex.realtime_commands.DUPLEX_REALTIME_CAPABILITIES.
input_audio_formats class-attribute instance-attribute ¶
input_audio_formats: frozenset[str] = field(
default_factory=lambda: frozenset(
REALTIME_INPUT_AUDIO_FORMATS
)
)
output_audio_formats class-attribute instance-attribute ¶
output_audio_formats: frozenset[str] = field(
default_factory=lambda: frozenset(
REALTIME_OUTPUT_AUDIO_FORMATS
)
)
validate_turn_detection class-attribute instance-attribute ¶
validate_turn_detection: TurnDetectionValidator | None = (
None
)
RealtimeProtocolError ¶
Bases: ValueError
A client payload could not be decoded into a valid Realtime intent.
code is the internal error code that :data:REALTIME_ERROR_TYPES_BY_CODE maps to an OpenAI error.type; event_id is the client event id the error answers. vllm_omni.engine.duplex.commands.DuplexCommandError is the duplex specialization, so a consumer catching either sees the same three attributes.
RealtimeSessionRejection dataclass ¶
Why a session object was refused: the internal error code and the message.
param names the offending session field where there is one, for the error.param slot of the OpenAI error envelope.
ResponseScopedEvent dataclass ¶
Bases: RealtimeEvent
Events addressed to one response (and usually one output item / content part).
apply_realtime_session_defaults ¶
apply_realtime_session_defaults(
defaults: RealtimeInputDefaults,
session_payload: Mapping[str, object],
) -> RealtimeInputDefaults
Derive the wire defaults a Realtime session object declares.
convert_input_audio_with_rate ¶
convert_input_audio_with_rate(
audio: object,
fmt: object,
*,
sample_rate_hz: int | float | None = None,
target_sample_rate_hz: int = 16000,
) -> tuple[object, object, int | float | None]
convert_output_audio ¶
convert_output_audio(
audio: str,
*,
source_fmt: str,
target_fmt: str,
source_sample_rate_hz: int | None = None,
target_sample_rate_hz: int | None = None,
) -> tuple[str, str, int | None]
copy_realtime_input_hints ¶
decode_audio_append ¶
decode_audio_append(
event: Mapping[str, object],
*,
defaults: RealtimeInputDefaults,
hints_source: Mapping[str, object] | None = None,
) -> RealtimeAudioAppend
Validate and pack one input append (audio and/or video frames).
Capability checks (required/optional modalities) run later on the session. Audio path converts to 16 kHz pcm_f32le.
hints_source is the enclosing payload when the audio arrives inside something larger than a bare append --- a conversation.item.create audio part, say --- so its hints apply unless the part overrides them.
Raises :class:RealtimeProtocolError for an unsupported format, undecodable audio or invalid camera frames.
encode_float32_mono_wav_base64 ¶
Package normalized mono float32 samples as a PCM16 WAV data payload.
input_audio_transcription_config ¶
input_audio_transcription_config(
session_payload: Mapping[str, object],
) -> dict[str, object] | None
input_looks_like_speech ¶
input_looks_like_speech(
event: Mapping[str, object],
*,
audio: object,
fmt: object,
overlap_silence_rms: float,
) -> bool
json_safe_realtime_payload ¶
normalize_conversation_item ¶
parse_realtime_audio_format ¶
realtime_audio_format_object ¶
realtime_audio_format_object(
fmt: object, *, sample_rate_hz: int | None = None
) -> dict[str, object]
realtime_max_output_tokens ¶
Normalize Realtime max output tokens ("inf" -> None).
realtime_overlap_fields ¶
resample_pcm16_mono ¶
text_chars_for_audio_ms_from_marks ¶
text_chars_for_audio_ms_from_marks(
audio_end_ms: int,
text_len: int,
marks: list[object],
*,
final_ms: object | None = None,
) -> int
truncate_realtime_item_content ¶
truncate_realtime_item_content(
item: dict[str, object],
*,
content_index: int,
audio_end_ms: int,
) -> None
validate_conversation_item_audio_formats ¶
validate_realtime_item_truncate ¶
validate_realtime_item_truncate(
item: Mapping[str, object],
*,
content_index: int,
audio_end_ms: int,
) -> str | None
validate_realtime_response_audio_formats ¶
validate_realtime_session_audio_formats ¶
validate_realtime_session_audio_formats(
session_payload: Mapping[str, object],
*,
input_audio_formats: Collection[str] | None = None,
output_audio_formats: Collection[str] | None = None,
) -> str | None
Reject a session object that declares an audio format we cannot serve.
The format sets default to everything the codec can decode; a consumer that serves a narrower set passes its own (see vllm_omni.protocol.realtime.capabilities).
validate_realtime_video_frames ¶
Validate omni-duplex camera frames on input_audio_buffer.append.
Wire contract matches the official MiniCPM-o duplex loop: one base base64 JPEG per ~1 s audio chunk, optionally followed by that unit's stacked composite tiling the sub-frames captured inside it (at most 2 images either way). A caller-supplied max_slice_nums is rejected rather than silently ignored: slicing is Stage 0's decision here, and Stage 0 already applies the official HD suggestion for a stacked unit (max_slice_nums=[2, 1]). The wire simply does not let the client choose it.
validate_session_payload ¶
validate_session_payload(
session_payload: Mapping[str, object],
*,
capabilities: RealtimeProtocolCapabilities,
) -> RealtimeSessionRejection | None
Check a session object against one consumer; None when it is acceptable.
Audio formats are checked before turn detection, so a session object that is wrong in both ways reports the format problem --- keep that order, it is what clients already see.