Skip to content

vllm_omni.entrypoints.utils

logger module-attribute

logger = init_logger(__name__)

PureDiffusionLauncherAdapter

vLLM launcher compatibility shim for pure-diffusion mode.

The upstream launcher's shutdown path reads app.state.engine_client.vllm_config.shutdown_timeout (vllm/entrypoints/launcher.py), but AsyncOmni.vllm_config returns None when the pipeline has no comprehension stage (pure diffusion), which crashes handle_shutdown with AttributeError and hangs server teardown (workers force-killed, spurious resource_tracker noise).

This adapter only overrides the vllm_config property with a minimal fallback carrying shutdown_timeout and forwards every other attribute to the wrapped engine client, so the pure-diffusion detection (get_vllm_config() still returns None) is unaffected.

vllm_config property

vllm_config: Any

StageConfigInputs dataclass

Normalized inputs shared by all stage-config loading paths.

deploy_config_path instance-attribute

deploy_config_path: str | None

kwargs instance-attribute

kwargs: dict[str, Any]

model instance-attribute

model: str

stage_overrides instance-attribute

stage_overrides: dict[str, dict[str, Any]] | None

strategy_config_path instance-attribute

strategy_config_path: str | None

trust_remote_code instance-attribute

trust_remote_code: bool | None

coerce_param_message_types

coerce_param_message_types(
    params: list[OmniSamplingParams], is_streaming: bool
)

Iterate over the sampling params and convert to the message types to DELTA messages, if streaming is enabled, or FINAL_ONLY if it's disabled, while respecting .skip_clone on the params.

This is needed to avoid emitting redundant multimodal data.

detect_pid_host

detect_pid_host() -> bool

filter_dataclass_kwargs

filter_dataclass_kwargs(cls: Any, kwargs: dict) -> dict

Filter kwargs to only include fields defined in the dataclass.

Parameters:

Name Type Description Default
cls Any

Dataclass type

required
kwargs dict

Keyword arguments to filter

required

Returns:

Type Description
dict

Filtered keyword arguments containing only valid dataclass fields

get_final_stage_id_for_e2e

get_final_stage_id_for_e2e(
    output_modalities: list[str] | None,
    default_modalities: list[str],
    stage_list: list,
) -> int

Get the final stage id for e2e.

Parameters:

Name Type Description Default
stage_list list

List of stage configurations

required

Returns:

Type Description
int

Final stage id for e2e

has_pid_host

has_pid_host() -> bool | None

Returns:

Type Description
bool | None

True -> very likely running with --pid=host (host PID namespace)

bool | None

False -> very likely isolated PID namespace (default)

bool | None

None -> cannot determine

in_container

in_container() -> bool

inject_omni_kv_config

inject_omni_kv_config(
    stage: Any,
    omni_conn_cfg: dict[str, Any],
    omni_from: str,
    omni_to: str,
) -> None

Inject connector configuration into stage engine arguments.

maybe_coerce_to_message_type

maybe_coerce_to_message_type(
    params: SamplingParams, is_streaming: bool
)

If this is a CUMULATIVE message, coerce it to DELTA if streaming, otherwise FINAL_ONLY.

parse_stage_overrides

parse_stage_overrides(
    value: Any,
) -> dict[str, dict[str, Any]] | None

Parse and validate the shape of per-stage JSON overrides.

prepare_stage_config_inputs

prepare_stage_config_inputs(
    model: str,
    kwargs: Mapping[str, Any],
    *,
    trust_remote_code: bool | None,
    snapshot_model: bool = False,
) -> StageConfigInputs

Normalize model/config arguments before resolving stage configs.

The standard API worker, multi-API parent, and headless entrypoint must apply the same legacy-key filtering, trust-remote-code precedence, and config-file extraction before calling :func:resolve_omni_config.