vllm_omni.outputs ¶
Modules:
| Name | Description |
|---|---|
duplex | |
mm_outputs | Multimodal output data structures for vLLM-Omni. |
multimodal_accumulation | |
output_metadata | |
output_modality | Output modality types for vLLM-Omni. |
output_processor | |
utils | Shared helpers for multimodal output handling. |
OmniConnectorOutput dataclass ¶
Communication results from Model Runner to Scheduler.
Carries transfer readiness signals so the Scheduler can make scheduling decisions without ever calling connector.put()/get() directly.
Attributes:
| Name | Type | Description |
|---|---|---|
chunk_ready_req_ids | set[str] | Request IDs with newly arrived chunks this cycle. |
chunk_finished_req_ids | set[str] | Request IDs whose final chunk has arrived. |
request_metadata | dict[str, dict[str, Any]] | Lightweight scheduling metadata keyed by request ID (e.g. next_stage_prompt_len, code_predictor_codes, left_context_size). Full payloads are owned by the Model Runner's local cache. |
kv_sent_req_ids | list[str] | Request IDs whose KV cache was successfully sent. |
stage_recv_req_ids | set[str] | Request IDs that received batch stage inputs. |
has_pending_kv_work | bool | True if the mixin has pending, active, or completed KV transfers that the scheduler should account for. |
OmniModelRunnerOutput dataclass ¶
Bases: ModelRunnerOutput
Model runner output for omni models.
Extends the base ModelRunnerOutput with support for multimodal outputs that may be produced by non-autoregressive stages.
Attributes:
| Name | Type | Description |
|---|---|---|
multimodal_outputs | list[dict[str, object]] | None | Optional per-request list of client-facing multimodal output dicts, indexed by req_index. |
inter_stage_outputs | list[dict[str, Any] | None] | None | Optional per-request list of inter-stage payload dicts for connector transport ( |
inter_stage_outputs class-attribute instance-attribute ¶
kv_extracted_req_ids class-attribute instance-attribute ¶
multimodal_outputs class-attribute instance-attribute ¶
omni_connector_output class-attribute instance-attribute ¶
omni_connector_output: OmniConnectorOutput | None = None
sampled_token_ids_materialized class-attribute instance-attribute ¶
sampled_token_ids_materialized: bool = False
with_kv_conn_output_only classmethod ¶
with_kv_conn_output_only(
kv_connector_output: Any,
) -> OmniModelRunnerOutput
OmniRequestOutput dataclass ¶
Bases: RequestOutput
Unified request output for both pipeline stages and diffusion models.
Extends vLLM's RequestOutput so that omni outputs can flow directly through vLLM serving codepaths (which expect prompt_token_ids, outputs, etc. as real attributes). The inherited fields store the LLM generation content; omni-specific fields store pipeline/diffusion extras.
Note: RequestOutput is a plain class (not a dataclass), so all of its attributes are redeclared below as dataclass fields with defaults — the dataclass-generated __init__ replaces RequestOutput.__init__ and must set them itself.
This class handles outputs from: 1. Multi-stage LLM pipelines (with stage_id, final_output_type, and the inherited RequestOutput fields carrying the stage's generation content) 2. Diffusion models (with images, prompt, metrics)
Attributes:
| Name | Type | Description |
|---|---|---|
request_id | str | Unique identifier for this request |
finished | bool | Whether generation is complete |
stage_id | int | None | Identifier of the stage that produced this output (pipeline mode) |
replica_id | int | None | Identifier of the stage replica that produced this output |
final_output_type | str | Type of output ("text", "image", "audio", "latents") |
images | list[Image] | List of generated PIL images (diffusion mode) |
prompt | OmniPromptType | None | The prompt used for generation |
latents | Tensor | None | Optional tensor of latent representations (diffusion mode) |
metrics | Any | Generation metrics. A plain dict for omni outputs; may carry vLLM's request stats object when copied from a raw RequestOutput. |
custom_output property writable ¶
Return custom output data from diffusion pipelines.
ec_transfer_params class-attribute instance-attribute ¶
encoder_prompt_token_ids class-attribute instance-attribute ¶
kv_transfer_params class-attribute instance-attribute ¶
multimodal_output property ¶
multimodal_output: Any
Return the multimodal output payload.
Checks completion outputs first (where multimodal_output is attached by AR stages), then the local _multimodal_output field.
Returns either a MultimodalPayload (Phase 3+) or a plain dict (legacy).
num_cache_creation_tokens class-attribute instance-attribute ¶
num_cache_creation_tokens: int | None = None
outputs class-attribute instance-attribute ¶
stage_durations class-attribute instance-attribute ¶
trajectory_log_probs class-attribute instance-attribute ¶
trajectory_timesteps class-attribute instance-attribute ¶
from_diffusion classmethod ¶
from_diffusion(
request_id: str,
images: list[Image],
prompt: OmniPromptType | None = None,
metrics: dict[str, Any] | None = None,
latents: Tensor | None = None,
trajectory_latents: Tensor | None = None,
trajectory_timesteps: Tensor | None = None,
trajectory_log_probs: Tensor | None = None,
trajectory_decoded: list | None = None,
multimodal_output: dict[str, Any] | None = None,
custom_output: dict[str, Any] | None = None,
final_output_type: str = "image",
stage_durations: dict[str, float] | None = None,
peak_memory_mb: float = 0.0,
finished: bool = True,
) -> OmniRequestOutput
Create output from diffusion model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
request_id | str | Request identifier | required |
images | list[Image] | Generated images | required |
prompt | OmniPromptType | None | The prompt used | None |
metrics | dict[str, Any] | None | Generation metrics | None |
latents | Tensor | None | Optional latent tensors | None |
trajectory_latents | Tensor | None | Optional stacked trajectory latent tensors | None |
trajectory_timesteps | Tensor | None | Optional stacked trajectory timestep tensors | None |
trajectory_log_probs | Tensor | None | Optional stacked trajectory log-probability tensors | None |
trajectory_decoded | list | None | Optional list of decoded trajectory images | None |
multimodal_output | dict[str, Any] | None | Optional multimodal output dict | None |
custom_output | dict[str, Any] | None | Optional custom output dict (e.g. prompt embeds) | None |
stage_durations | dict[str, float] | None | Optional stage durations (execution time of each stage) dict | None |
peak_memory_mb | float | Peak memory usage in MB | 0.0 |
Returns:
| Type | Description |
|---|---|
OmniRequestOutput | OmniRequestOutput configured for diffusion mode |
from_error classmethod ¶
from_error(
request_id: str,
error_message: str,
*,
status_code: int | None = None,
error_type: str | None = None,
) -> OmniRequestOutput
Create a terminal error output.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
request_id | str | Request identifier | required |
error_message | str | Human-readable error description | required |
Returns:
| Type | Description |
|---|---|
OmniRequestOutput | OmniRequestOutput with |
from_stage_output classmethod ¶
from_stage_output(
source: RequestOutput, **kwargs: Any
) -> OmniRequestOutput
Create an OmniRequestOutput from a stage's raw output.
Copies generation content (outputs, prompt, prompt_token_ids, finished, images, latents, etc.) from source onto the returned object. source may be a vLLM RequestOutput, another OmniRequestOutput (which inherits from RequestOutput).
This is the preferred way to construct an OmniRequestOutput that wraps a stage result.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source | RequestOutput | The stage output whose content is copied onto the new object. | required |
**kwargs | Any | Passed through to the dataclass constructor ( | {} |
Returns:
| Type | Description |
|---|---|
OmniRequestOutput | A new |