API Server¶
vLLM-Omni exposes OpenAI-compatible HTTP APIs plus vLLM-Omni extensions for image, audio, video, realtime, and robot-policy workloads. Use this page to choose an endpoint; follow the linked reference page for model-specific fields, wire protocols, and response details.
Each server instance hosts one model. An endpoint is usable only when the loaded model supports that task.
This guide focuses on vLLM-Omni's primary task endpoints. Compatible models may also expose standard routes inherited from vLLM, such as POST /v1/completions; those routes remain model- and task-dependent.
Start and Verify the Server¶
Start a model with the shared serving command:
Some models require additional flags or a deployment configuration. Use the model-specific guide when one is provided.
After the server starts, check its health and served model name:
export VLLM_OMNI_BASE_URL=http://localhost:8091
curl "$VLLM_OMNI_BASE_URL/health"
curl "$VLLM_OMNI_BASE_URL/v1/models" | jq .
Use http://localhost:8091/v1 as the base_url for an OpenAI SDK client. If the server was started with --api-key, include an Authorization header with Bearer <api-key> in requests.
Choose a Core HTTP Endpoint¶
Prefer a task-specific endpoint when one matches your workload. Use Chat Completions for conversational or heterogeneous omni pipelines, rather than as the default wrapper for every generation task.
| Task | Endpoint | Request | Response | Details |
|---|---|---|---|---|
| Conversation or multimodal understanding/generation | POST /v1/chat/completions | JSON | OpenAI-style JSON or SSE | Chat Completions |
| Text-to-speech | POST /v1/audio/speech | JSON | Audio bytes or SSE | Speech |
| Sound, music, or ambient audio generation | POST /v1/audio/generate | JSON | Audio bytes | Audio Generation |
| Text-to-image | POST /v1/images/generations | JSON | Base64 JSON or an image file | Image Generation |
| Image editing | POST /v1/images/edits | Multipart form | Base64 JSON or SSE | Image Edit |
| Video generation (recommended) | POST /v1/videos | Multipart form | Asynchronous job JSON | Videos |
| Video generation (blocking) | POST /v1/videos/sync | Multipart form | Video bytes | Videos |
POST /v1/videos/sync is intended for tests and simple scripts. Use the asynchronous job API for production requests so clients can poll, download, and delete generated videos independently.
Minimal Requests¶
The examples below assume a compatible model is already running. TTS voices, media inputs, generation controls, and output capabilities vary by model.
curl "$VLLM_OMNI_BASE_URL/v1/images/edits" \
-F "[email protected]" \
-F "prompt=Turn the background into a snowy mountain" \
| jq -r '.data[0].b64_json' \
| base64 --decode > edited.png
Choose a Streaming or Realtime Endpoint¶
These WebSocket APIs have different event schemas and cannot be used interchangeably.
| Workload | Endpoint | Interaction model | Details |
|---|---|---|---|
| Incremental text input for speech synthesis | WS /v1/audio/speech/stream | Send text events and receive audio | Streaming Text to Speech |
| Live video understanding | WS /v1/video/chat/stream | Send video frames and receive text/audio | Streaming Video Input |
| Turn-based realtime audio | WS /v1/realtime | Stream one audio input and receive transcript/audio events | Realtime Audio |
| Continuous speech-to-speech interaction | WS /v1/realtime?duplex=1 (alias WS /v1/duplex) | Listen and speak concurrently with session control | Full Duplex |
| Generated video chunks | WS /v1/realtime/video | Start a diffusion request and receive fragmented MP4 | Streaming Video Output |
| Robot policy inference | WS /v1/realtime/robot/openpi | Send MessagePack observations and receive action arrays | OpenPI Robot Policy |
All six routes are model- or configuration-dependent. In particular, /v1/realtime is full duplex when the server is a duplex server: the model's pipeline declares a duplex_plugin and its deploy configuration sets session_mode: duplex (with session_mode: turn the same model boots the ordinary turn-based serving stack instead). On a duplex server a stock Realtime client needs no vendor query parameter; duplex=0 selects the turn-based Realtime handler, which a duplex server does not mount, so that connection is refused with Realtime API is not available. Such a server serves the websocket route plus POST /v1/chat/completions, /v1/models and /health, and no other turn-based HTTP route. The chat route is the ordinary chat service running on the duplex engine: a request is a turn-based generation on the same stages, served alongside the live websocket sessions; it opens no duplex session and holds no duplex_session.max_sessions slot -- see Full Duplex. On a server that is not duplex, an explicit ?duplex=1 is refused rather than answered by the turn-based handler, so a client that requested duplex never silently gets the other protocol.
Related Endpoints¶
| Purpose | Endpoints | Reference |
|---|---|---|
| Discovery and readiness | GET /health, GET /v1/models | This page |
| Batched conversations | POST /v1/chat/completions/batch | Batch requests |
| Batched speech | POST /v1/audio/speech/batch | Batch speech generation |
| TTS voice management | GET/POST /v1/audio/voices, DELETE /v1/audio/voices/{name} | Voices |
| Video job lifecycle | GET /v1/videos, GET/DELETE /v1/videos/{video_id}, GET /v1/videos/{video_id}/content | Video endpoints |
| Release and restore stage memory | POST /v1/omni/sleep, POST /v1/omni/wakeup | Sleep Mode |
Standalone Experimental Servers¶
JoyVL provides a stateful interaction orchestrator in front of another model server. It is a separate process with its own routes; they are not additional paths on every vllm serve --omni instance. See Standalone Experimental Servers before deploying it.
Compatibility Notes¶
- A registered path does not mean every model supports it. Unsupported tasks return an OpenAI-style error or a WebSocket
errorevent. - Use JSON for Chat, Speech, Audio Generation, and Image Generation. Use
multipart/form-datawhen uploading images, videos, or other files. - Generation can take minutes for large diffusion models. Configure client timeouts accordingly; prefer asynchronous video jobs over long blocking requests.
- Model-specific parameters are vLLM-Omni extensions. Check the endpoint reference before assuming an OpenAI SDK exposes them directly.