Skip to content

API Server

vLLM-Omni exposes OpenAI-compatible HTTP APIs plus vLLM-Omni extensions for image, audio, video, realtime, and robot-policy workloads. Use this page to choose an endpoint; follow the linked reference page for model-specific fields, wire protocols, and response details.

Each server instance hosts one model. An endpoint is usable only when the loaded model supports that task.

This guide focuses on vLLM-Omni's primary task endpoints. Compatible models may also expose standard routes inherited from vLLM, such as POST /v1/completions; those routes remain model- and task-dependent.

Start and Verify the Server

Start a model with the shared serving command:

vllm serve <model> --omni --port 8091

Some models require additional flags or a deployment configuration. Use the model-specific guide when one is provided.

After the server starts, check its health and served model name:

export VLLM_OMNI_BASE_URL=http://localhost:8091

curl "$VLLM_OMNI_BASE_URL/health"
curl "$VLLM_OMNI_BASE_URL/v1/models" | jq .

Use http://localhost:8091/v1 as the base_url for an OpenAI SDK client. If the server was started with --api-key, include an Authorization header with Bearer <api-key> in requests.

Choose a Core HTTP Endpoint

Prefer a task-specific endpoint when one matches your workload. Use Chat Completions for conversational or heterogeneous omni pipelines, rather than as the default wrapper for every generation task.

Task Endpoint Request Response Details
Conversation or multimodal understanding/generation POST /v1/chat/completions JSON OpenAI-style JSON or SSE Chat Completions
Text-to-speech POST /v1/audio/speech JSON Audio bytes or SSE Speech
Sound, music, or ambient audio generation POST /v1/audio/generate JSON Audio bytes Audio Generation
Text-to-image POST /v1/images/generations JSON Base64 JSON or an image file Image Generation
Image editing POST /v1/images/edits Multipart form Base64 JSON or SSE Image Edit
Video generation (recommended) POST /v1/videos Multipart form Asynchronous job JSON Videos
Video generation (blocking) POST /v1/videos/sync Multipart form Video bytes Videos

POST /v1/videos/sync is intended for tests and simple scripts. Use the asynchronous job API for production requests so clients can poll, download, and delete generated videos independently.

Minimal Requests

The examples below assume a compatible model is already running. TTS voices, media inputs, generation controls, and output capabilities vary by model.

curl "$VLLM_OMNI_BASE_URL/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Describe vLLM-Omni briefly."}],
    "modalities": ["text"]
  }'
curl "$VLLM_OMNI_BASE_URL/v1/audio/speech" \
  -H "Content-Type: application/json" \
  -d '{
    "input": "Hello from vLLM-Omni.",
    "voice": "vivian"
  }' \
  --output speech.wav
curl "$VLLM_OMNI_BASE_URL/v1/audio/generate" \
  -H "Content-Type: application/json" \
  -d '{
    "input": "Ocean waves on a quiet beach",
    "audio_length": 5.0
  }' \
  --output audio.wav
curl "$VLLM_OMNI_BASE_URL/v1/images/generations" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "A small robot reading beside a window",
    "size": "1024x1024",
    "response_format": "file"
  }' \
  --output image.png
curl "$VLLM_OMNI_BASE_URL/v1/images/edits" \
  -F "[email protected]" \
  -F "prompt=Turn the background into a snowy mountain" \
  | jq -r '.data[0].b64_json' \
  | base64 --decode > edited.png
curl "$VLLM_OMNI_BASE_URL/v1/videos" \
  -F "prompt=A cinematic tracking shot of a mountain lake at sunrise" \
  -F "size=1280x720" \
  -F "seconds=5"

The response contains a video job id. Poll and download it with:

curl "$VLLM_OMNI_BASE_URL/v1/videos/<video_id>"
curl -L "$VLLM_OMNI_BASE_URL/v1/videos/<video_id>/content" --output video.mp4

Choose a Streaming or Realtime Endpoint

These WebSocket APIs have different event schemas and cannot be used interchangeably.

Workload Endpoint Interaction model Details
Incremental text input for speech synthesis WS /v1/audio/speech/stream Send text events and receive audio Streaming Text to Speech
Live video understanding WS /v1/video/chat/stream Send video frames and receive text/audio Streaming Video Input
Turn-based realtime audio WS /v1/realtime Stream one audio input and receive transcript/audio events Realtime Audio
Continuous speech-to-speech interaction WS /v1/realtime?duplex=1 (alias WS /v1/duplex) Listen and speak concurrently with session control Full Duplex
Generated video chunks WS /v1/realtime/video Start a diffusion request and receive fragmented MP4 Streaming Video Output
Robot policy inference WS /v1/realtime/robot/openpi Send MessagePack observations and receive action arrays OpenPI Robot Policy

All six routes are model- or configuration-dependent. In particular, /v1/realtime is full duplex when the server is a duplex server: the model's pipeline declares a duplex_plugin and its deploy configuration sets session_mode: duplex (with session_mode: turn the same model boots the ordinary turn-based serving stack instead). On a duplex server a stock Realtime client needs no vendor query parameter; duplex=0 selects the turn-based Realtime handler, which a duplex server does not mount, so that connection is refused with Realtime API is not available. Such a server serves the websocket route plus POST /v1/chat/completions, /v1/models and /health, and no other turn-based HTTP route. The chat route is the ordinary chat service running on the duplex engine: a request is a turn-based generation on the same stages, served alongside the live websocket sessions; it opens no duplex session and holds no duplex_session.max_sessions slot -- see Full Duplex. On a server that is not duplex, an explicit ?duplex=1 is refused rather than answered by the turn-based handler, so a client that requested duplex never silently gets the other protocol.

Purpose Endpoints Reference
Discovery and readiness GET /health, GET /v1/models This page
Batched conversations POST /v1/chat/completions/batch Batch requests
Batched speech POST /v1/audio/speech/batch Batch speech generation
TTS voice management GET/POST /v1/audio/voices, DELETE /v1/audio/voices/{name} Voices
Video job lifecycle GET /v1/videos, GET/DELETE /v1/videos/{video_id}, GET /v1/videos/{video_id}/content Video endpoints
Release and restore stage memory POST /v1/omni/sleep, POST /v1/omni/wakeup Sleep Mode

Standalone Experimental Servers

JoyVL provides a stateful interaction orchestrator in front of another model server. It is a separate process with its own routes; they are not additional paths on every vllm serve --omni instance. See Standalone Experimental Servers before deploying it.

Compatibility Notes

  • A registered path does not mean every model supports it. Unsupported tasks return an OpenAI-style error or a WebSocket error event.
  • Use JSON for Chat, Speech, Audio Generation, and Image Generation. Use multipart/form-data when uploading images, videos, or other files.
  • Generation can take minutes for large diffusion models. Configure client timeouts accordingly; prefer asynchronous video jobs over long blocking requests.
  • Model-specific parameters are vLLM-Omni extensions. Check the endpoint reference before assuming an OpenAI SDK exposes them directly.