vllm_omni.entrypoints.omni_base ¶
OutputMessageHandleResult module-attribute ¶
OutputMessageHandleResult = (
tuple[Literal[True], None, None, None]
| tuple[Literal[False], str, int, ClientRequestState]
)
OmniBase ¶
Bases: PDDisaggregationMixin
Shared runtime foundation for AsyncOmni and Omni.
default_sampling_params_list instance-attribute ¶
engine instance-attribute ¶
engine = self._create_engine(
model=model,
init_timeout=init_timeout,
stage_init_timeout=stage_init_timeout,
transfer_emitter=self.transfer_metrics,
prom_metrics=self.prom_metrics,
log_stats=log_stats,
**kwargs,
)
errored property ¶
errored: bool
Whether the engine is in a process-fatal error state.
True only when the orchestrator thread is dead. Per-stage liveness is deliberately excluded: the OpenAI serving paths precheck errored before request routing, so including it would reject requests that do not touch the dead stage (e.g. text-only chat when the talker stage is down). A fully dead stage instead fails only the requests routed through it (dispatch guards) and flips readiness to 503 via :meth:check_health (per-replica fault isolation, #4285).
mod_metrics instance-attribute ¶
mod_metrics = OmniModalityMetrics(
model_name=model, log_stats=log_stats
)
prom_metrics instance-attribute ¶
prom_metrics = OmniPrometheusMetrics(
model_name=model, log_stats=log_stats
)
sampling_constraints_list instance-attribute ¶
stage_configs property ¶
stage_configs: list
Expose engine stage configs for PD disaggregation detection and validation.
transfer_metrics instance-attribute ¶
transfer_metrics = OmniTransferMetrics(
model_name=model, log_stats=log_stats
)
tts_batch_max_items instance-attribute ¶
tts_batch_max_items: int = kwargs.pop(
"tts_batch_max_items", 32
)
from_cli_args classmethod ¶
from_cli_args(
args: TrackingNamespace, model: str | None = None
) -> OmniBase
Build from a TrackingNamespace parsed by TrackingArgumentParser. Only args that are explicitly passed to parse_args are forwarded.
resolve_sampling_params_list ¶
resolve_sampling_params_list(
sampling_params_list: Sequence[Any] | Any | None,
allow_delta_coercion: bool = False,
) -> Sequence[Any]
Resolve request parameters; pipeline sampling constraints override caller values.
start_profile ¶
Start profiling specified stages.
Uses vLLM-compatible profile(is_start=True, profile_prefix) interface.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
profile_prefix | str | None | Optional prefix for the trace file names. | None |
stages | list[int] | None | List of stage IDs to profile. If None, profiles all stages. | None |
Returns:
| Type | Description |
|---|---|
list[Any] | List of results from each stage. |