Skip to content

vllm_omni.protocol.realtime

The OpenAI Realtime wire codec, shared by every vLLM-Omni Realtime surface.

What this is

The model-agnostic half of WS /v1/realtime: how a client event is parsed, which audio formats and session fields are accepted, what a conversation item means, how audio is decoded and resampled, and the shape of the error envelope (the codes inside it are a consumer's own --- ours are in vllm_omni.protocol.duplex.errors). It is a library of pure functions and value objects.

What it is not

It owns no session. It does not decide when a response starts, what a model is prompted with, whether a turn is over, or where committed audio goes. Those are runtime decisions, and after PR #7413 a duplex session's runtime decisions all belong to vllm_omni.engine.duplex --- the session runner, the typed DuplexCommand / DuplexEvent boundary and the DuplexModelPlugin seam. This package sits under that: the duplex engine binds the codec to its own command and event vocabulary in vllm_omni.engine.duplex.realtime_commands and vllm_omni.engine.duplex.realtime_events.

Why it is separate (RFC #6592, P0a)

So that a second Realtime surface --- a Qwen3-Omni GA profile, say --- can reuse the parsing, validation, format negotiation and error envelope without adopting the duplex session control plane (lease manager, commit policy, model channel), and without a second copy of the codec drifting away from this one. A consumer states what it can serve through :class:~vllm_omni.protocol.realtime.capabilities.RealtimeProtocolCapabilities and keeps its own session state.

Known duplex-shaped defaults

This package is model-agnostic in its dependencies, not yet in every default. Two constants still carry the assumptions of its first consumer and a second one will have to parameterize them:

  • :func:~vllm_omni.protocol.realtime.audio.convert_input_audio_with_rate resamples to 16 kHz, which is MiniCPM-o's feature-extractor rate. A GA client sending 24 kHz is converted down rather than served at its own rate.
  • :func:~vllm_omni.protocol.realtime.items.validate_realtime_video_frames encodes the MiniCPM-o duplex camera contract (at most two frames per append, max_slice_nums not selectable on the wire).

Both are behaviour this extraction deliberately did not change (RFC #6592 P0a is behaviour-preserving); making them per-consumer is P0b/P1 work.

Dependency rule

Nothing here may import vllm_omni.engine, vllm_omni.entrypoints, vllm_omni.model_executor, vllm_omni.worker or vllm_omni.clients; tests/protocol/realtime/test_protocol_import_boundary.py asserts it.

Modules:

Name Description
audio
audio_input

Decoding one input_audio_buffer.append.

capabilities

What a Realtime consumer can serve, and the session check that reads it.

commands

OpenAI Realtime client events as typed commands.

errors

The OpenAI Realtime error envelope.

events

OpenAI Realtime server events: typed classes and their wire rendering.

formats

Realtime audio-format negotiation: what a client may declare and how it is read.

items

Realtime conversation items: shape, truncation and the transcript inside one.

session

Reading the Realtime session object.

MAX_INPUT_SAMPLE_RATE_HZ module-attribute

MAX_INPUT_SAMPLE_RATE_HZ = 192000

MIN_INPUT_SAMPLE_RATE_HZ module-attribute

MIN_INPUT_SAMPLE_RATE_HZ = 8000

REALTIME_INPUT_AUDIO_FORMATS module-attribute

REALTIME_INPUT_AUDIO_FORMATS = {
    "pcm16",
    "pcm_s16le",
    "s16le",
    "pcm_f32le",
    "g711_ulaw",
    "g711_alaw",
}

REALTIME_INPUT_HINT_KEYS module-attribute

REALTIME_INPUT_HINT_KEYS = (
    "duration_ms",
    "audio_duration_ms",
    "audio_start_ms",
    "audio_end_ms",
    "is_speech",
    "speech",
    "speech_probability",
    "vad",
    "overlap_action",
    "overlap",
    "force_barge_in",
    "force_listen",
    "text",
    "transcript",
)

REALTIME_OUTPUT_AUDIO_FORMATS module-attribute

REALTIME_OUTPUT_AUDIO_FORMATS = {
    "pcm16",
    "pcm_s16le",
    "s16le",
    "wav",
    "pcm",
    "g711_ulaw",
    "g711_alaw",
}

RealtimeAudioAppend dataclass

One decoded input_audio_buffer.append.

audio is raw bytes in format at sample_rate_hz --- after :func:decode_audio_append that is 16 kHz pcm_f32le for every input format the codec converts.

audio instance-attribute

audio: bytes

audio_end_ms class-attribute instance-attribute

audio_end_ms: int | None = None

duration_ms class-attribute instance-attribute

duration_ms: int | None = None

event_id class-attribute instance-attribute

event_id: str | None = None

format instance-attribute

format: str

hints class-attribute instance-attribute

hints: dict[str, object] = field(default_factory=dict)

is_speech class-attribute instance-attribute

is_speech: bool | None = None

sample_rate_hz class-attribute instance-attribute

sample_rate_hz: int | None = None

video_frames class-attribute instance-attribute

video_frames: tuple[str, ...] = ()

RealtimeCommand dataclass

One decoded OpenAI Realtime client event.

wire_type is the event type the client sent; the typed fields are what survived decoding, so a consumer reads them instead of re-parsing JSON.

There is deliberately no rendering method here. A server receives commands, it does not emit them, and how a runtime represents one internally is its own business --- the duplex engine renders its mailbox dictionary in vllm_omni.engine.duplex.commands.

event_id class-attribute instance-attribute

event_id: str | None = None

wire_type class-attribute

wire_type: str = ''

RealtimeEvent dataclass

Base of every public session event.

audio property

audio: bytes | None

epoch class-attribute instance-attribute

epoch: int | None = None

event_id class-attribute instance-attribute

event_id: str = field(default_factory=new_event_id)

is_terminal property

is_terminal: bool

item_id property

item_id: str | None

optional_wire_fields class-attribute

optional_wire_fields: frozenset[str] = frozenset()

response_id property

response_id: str | None

session_id class-attribute instance-attribute

session_id: str = ''

text property

text: str | None

type property

type: str

wire_type class-attribute

wire_type: str = ''

to_realtime

to_realtime() -> dict[str, object]

The Realtime wire JSON object for this event (derived from the fields).

RealtimeInputDefaults dataclass

Session-level wire defaults an append may omit (derived from the session object).

input_audio_format class-attribute instance-attribute

input_audio_format: str = 'pcm16'

input_sample_rate_hz class-attribute instance-attribute

input_sample_rate_hz: int = 16000

output_audio_format class-attribute instance-attribute

output_audio_format: str = 'pcm16'

output_sample_rate_hz class-attribute instance-attribute

output_sample_rate_hz: int | None = None

overlap_silence_rms class-attribute instance-attribute

overlap_silence_rms: float = 0.003

with_session_payload

with_session_payload(
    session_payload: Mapping[str, object],
) -> RealtimeInputDefaults

Return defaults updated from a Realtime session object (session.update).

RealtimeProtocolCapabilities dataclass

One consumer's answer to what a session object may ask for.

The defaults are permissive, not restrictive: every format the codec can decode, and no turn-detection validation at all. A consumer that leaves validate_turn_detection at None therefore accepts whatever turn_detection the session object carries --- including values it cannot actually serve, such as semantic_vad. A consumer that supports only some turn-detection modes (or none) must supply its own validator; see vllm_omni.engine.duplex.realtime_commands.DUPLEX_REALTIME_CAPABILITIES.

input_audio_formats class-attribute instance-attribute

input_audio_formats: frozenset[str] = field(
    default_factory=lambda: frozenset(
        REALTIME_INPUT_AUDIO_FORMATS
    )
)

output_audio_formats class-attribute instance-attribute

output_audio_formats: frozenset[str] = field(
    default_factory=lambda: frozenset(
        REALTIME_OUTPUT_AUDIO_FORMATS
    )
)

validate_turn_detection class-attribute instance-attribute

validate_turn_detection: TurnDetectionValidator | None = (
    None
)

RealtimeProtocolError

Bases: ValueError

A client payload could not be decoded into a valid Realtime intent.

code is the internal error code that :data:REALTIME_ERROR_TYPES_BY_CODE maps to an OpenAI error.type; event_id is the client event id the error answers. vllm_omni.engine.duplex.commands.DuplexCommandError is the duplex specialization, so a consumer catching either sees the same three attributes.

code instance-attribute

code = code

event_id instance-attribute

event_id = event_id

RealtimeSessionRejection dataclass

Why a session object was refused: the internal error code and the message.

param names the offending session field where there is one, for the error.param slot of the OpenAI error envelope.

code instance-attribute

code: str

message instance-attribute

message: str

param class-attribute instance-attribute

param: str | None = None

ResponseScopedEvent dataclass

Bases: RealtimeEvent

Events addressed to one response (and usually one output item / content part).

content_index class-attribute instance-attribute

content_index: int = 0

item_id class-attribute instance-attribute

item_id: str | None = None

output_index class-attribute instance-attribute

output_index: int = 0

response_id class-attribute instance-attribute

response_id: str | None = None

apply_realtime_session_defaults

apply_realtime_session_defaults(
    defaults: RealtimeInputDefaults,
    session_payload: Mapping[str, object],
) -> RealtimeInputDefaults

Derive the wire defaults a Realtime session object declares.

convert_input_audio_with_rate

convert_input_audio_with_rate(
    audio: object,
    fmt: object,
    *,
    sample_rate_hz: int | float | None = None,
    target_sample_rate_hz: int = 16000,
) -> tuple[object, object, int | float | None]

convert_output_audio

convert_output_audio(
    audio: str,
    *,
    source_fmt: str,
    target_fmt: str,
    source_sample_rate_hz: int | None = None,
    target_sample_rate_hz: int | None = None,
) -> tuple[str, str, int | None]

copy_realtime_input_hints

copy_realtime_input_hints(
    source: Mapping[str, object], target: dict[str, object]
) -> None

decode_audio_append

decode_audio_append(
    event: Mapping[str, object],
    *,
    defaults: RealtimeInputDefaults,
    hints_source: Mapping[str, object] | None = None,
) -> RealtimeAudioAppend

Validate and pack one input append (audio and/or video frames).

Capability checks (required/optional modalities) run later on the session. Audio path converts to 16 kHz pcm_f32le.

hints_source is the enclosing payload when the audio arrives inside something larger than a bare append --- a conversation.item.create audio part, say --- so its hints apply unless the part overrides them.

Raises :class:RealtimeProtocolError for an unsupported format, undecodable audio or invalid camera frames.

encode_float32_mono_wav_base64

encode_float32_mono_wav_base64(
    samples: ndarray, *, sample_rate_hz: int
) -> str

Package normalized mono float32 samples as a PCM16 WAV data payload.

input_audio_transcription_config

input_audio_transcription_config(
    session_payload: Mapping[str, object],
) -> dict[str, object] | None

input_explicitly_non_speech

input_explicitly_non_speech(
    event: Mapping[str, object],
) -> bool

input_looks_like_speech

input_looks_like_speech(
    event: Mapping[str, object],
    *,
    audio: object,
    fmt: object,
    overlap_silence_rms: float,
) -> bool

input_transcript_from_item

input_transcript_from_item(
    item: Mapping[str, object],
) -> str

is_supported_realtime_input_format

is_supported_realtime_input_format(fmt: object) -> bool

json_safe_realtime_payload

json_safe_realtime_payload(
    payload: Mapping[str, object],
) -> dict[str, object]

new_event_id

new_event_id() -> str

normalize_conversation_item

normalize_conversation_item(
    item: Mapping[str, object],
) -> dict[str, object]

parse_realtime_audio_format

parse_realtime_audio_format(
    raw_format: object,
) -> tuple[object, int | None]

realtime_audio_format_object

realtime_audio_format_object(
    fmt: object, *, sample_rate_hz: int | None = None
) -> dict[str, object]

realtime_max_output_tokens

realtime_max_output_tokens(value: object) -> int | None

Normalize Realtime max output tokens ("inf" -> None).

realtime_output_format

realtime_output_format(duplex_format: object) -> str

realtime_overlap_fields

realtime_overlap_fields(
    session_payload: Mapping[str, object],
) -> dict[str, object]

resample_pcm16_mono

resample_pcm16_mono(
    raw: bytes, *, source_rate_hz: int, target_rate_hz: int
) -> bytes

text_chars_for_audio_ms_from_marks

text_chars_for_audio_ms_from_marks(
    audio_end_ms: int,
    text_len: int,
    marks: list[object],
    *,
    final_ms: object | None = None,
) -> int

truncate_realtime_item_content

truncate_realtime_item_content(
    item: dict[str, object],
    *,
    content_index: int,
    audio_end_ms: int,
) -> None

validate_conversation_item_audio_formats

validate_conversation_item_audio_formats(
    item: object,
) -> str | None

validate_input_sample_rate_hz

validate_input_sample_rate_hz(
    sample_rate_hz: object,
) -> int

validate_realtime_item_truncate

validate_realtime_item_truncate(
    item: Mapping[str, object],
    *,
    content_index: int,
    audio_end_ms: int,
) -> str | None

validate_realtime_response_audio_formats

validate_realtime_response_audio_formats(
    response_payload: Mapping[str, object],
) -> str | None

validate_realtime_session_audio_formats

validate_realtime_session_audio_formats(
    session_payload: Mapping[str, object],
    *,
    input_audio_formats: Collection[str] | None = None,
    output_audio_formats: Collection[str] | None = None,
) -> str | None

Reject a session object that declares an audio format we cannot serve.

The format sets default to everything the codec can decode; a consumer that serves a narrower set passes its own (see vllm_omni.protocol.realtime.capabilities).

validate_realtime_video_frames

validate_realtime_video_frames(
    video_frames: object, max_slice_nums: object
) -> str | None

Validate omni-duplex camera frames on input_audio_buffer.append.

Wire contract matches the official MiniCPM-o duplex loop: one base base64 JPEG per ~1 s audio chunk, optionally followed by that unit's stacked composite tiling the sub-frames captured inside it (at most 2 images either way). A caller-supplied max_slice_nums is rejected rather than silently ignored: slicing is Stage 0's decision here, and Stage 0 already applies the official HD suggestion for a stacked unit (max_slice_nums=[2, 1]). The wire simply does not let the client choose it.

validate_session_payload

validate_session_payload(
    session_payload: Mapping[str, object],
    *,
    capabilities: RealtimeProtocolCapabilities,
) -> RealtimeSessionRejection | None

Check a session object against one consumer; None when it is acceptable.

Audio formats are checked before turn detection, so a session object that is wrong in both ways reports the format problem --- keep that order, it is what clients already see.

wav_payload_to_pcm16

wav_payload_to_pcm16(
    raw: bytes,
) -> tuple[bytes | None, int | None]

wire_value

wire_value(value: object) -> object