vllm_omni.worker ¶
Modules:
| Name | Description |
|---|---|
base | Base worker class for vLLM-Omni with device-level GPU memory profiling. |
gpu_ar_model_runner | AR GPU Model Runner for vLLM-Omni. |
gpu_ar_worker | |
gpu_generation_model_runner | Code2Wav GPU Model Runner for vLLM-Omni. |
gpu_generation_worker | |
gpu_memory_utils | NVML-based per-process GPU memory utilities. |
gpu_model_runner | |
memory_utils | GPU memory utilities for vLLM Omni workers. |
mixins | |
omni_connector_model_runner_mixin | Public model-runner interface for Omni connector transport. |
omni_connector_validation | Validation for Omni connector support on selected LLM workers. |
output | Worker output helpers. |
payload_span | Helpers for explicit thinker decode span metadata. |
runner_assisted_metadata | |
sampling_utils | Sampling helpers shared by the GPU and NPU AR model runners. |
sparse_audio | Sparse-audio output protocol: marker classification and payload routing. |