MiniCPM-o 4.5: Offline inference¶
Source https://github.com/vllm-project/vllm-omni/tree/main/examples/offline_inference/minicpmo.
Two-stage pipeline: thinker (multimodal understanding) → talker + Token2Wav (24 kHz speech). Deploy config auto-loads from vllm_omni/deploy/minicpmo_4_5.yaml (2-GPU default).
Setup¶
- Install talker deps:
pip install 'vllm-omni[minicpmo]'(orstepaudio2-minicpmo) --trust-remote-codeis always passed byend2end.py(required for MiniCPMO)- See stage configuration docs for memory tuning
Run examples¶
Single prompt¶
Multiple prompts¶
Thinker tensor parallel (3-GPU)¶
Modality control¶
Text-only (skip talker / no <|tts_bos|>):
Text + speech (default — appends <|tts_bos|>):
Local media files¶
python end2end.py --query-type use_video --video-path /path/to/video.mp4
python end2end.py --query-type use_image --image-path /path/to/image.jpg
python end2end.py --query-type use_audio --audio-path /path/to/audio.wav
python end2end.py --query-type use_mixed_modalities \
--video-path /path/to/video.mp4 \
--image-path /path/to/image.jpg \
--audio-path /path/to/audio.wav
Supported --query-type values:
| Query type | Inputs |
|---|---|
text | Text only |
use_image | Image + text |
use_audio | Audio + text |
use_video | Video + text |
use_multi_audios | Two audio clips |
use_mixed_modalities | Audio + image + video |
Custom deploy config¶
python end2end.py --query-type text \
--deploy-config /path/to/vllm_omni/deploy/minicpmo_4_5_8x4090.yaml
Notes¶
- Speech requires
<|tts_bos|>on the assistant prefix (offline equivalent of onlinechat_template_kwargs.use_tts_template=true). Without it, the talker gets an empty TTS span. - Output WAV is 24 kHz mono.
- Placeholders in the prompt are MiniCPM-style:
(<image>./</image>),(<audio>./</audio>),(<video>./</video>). - Default layout needs 2 GPUs. Async chunking is off in the bundled YAMLs.
Online serving¶
See examples/online_serving/minicpmo/ and the recipe recipes/OpenBMB/MiniCPM-o-4_5.md.
Example materials¶
end2end.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/offline_inference/minicpmo/end2end.py.
run_multiple_prompts.sh
run_single_prompt.sh
run_single_prompt_tp.sh
#!/usr/bin/env bash
# Single-prompt offline run with thinker tensor-parallel (3-GPU layout).
# Thinker on GPU 0,1 (TP=2); talker + Token2Wav on GPU 2.
set -euo pipefail
cd "$(dirname "$0")"
REPO_ROOT="$(cd ../../.. && pwd)"
python end2end.py --output-wav output_audio \
--query-type use_audio \
--deploy-config "${REPO_ROOT}/vllm_omni/deploy/minicpmo_4_5_3gpu.yaml" \
--stage-init-timeout 300