Additional Configuration¶
Additional configuration is a mechanism provided by vLLM to allow plugins to control internal behavior by themselves. VLLM Ascend uses this mechanism to make the project more flexible.
Migration Guide¶
Starting from PR #9064, vLLM Ascend is migrating supported environment variables to --additional-config.
Important Notice¶
- Current Support: Both environment variables and
--additional-configare supported during the transition period - Recommendation: Use
--additional-configfor new deployments and migrate existing configurations - Future Plan: Environment variables will be removed in a future release; only
--additional-configwill be supported
Quick Reference¶
| Environment Variable | Config Key | Type Conversion |
|---|---|---|
VLLM_ASCEND_BALANCE_SCHEDULING |
scheduler_config.enable_balance_scheduling |
"1" → true, "0" → false |
MSMONITOR_USE_DAEMON |
msmonitor_use_daemon |
"1" → true, "0" → false |
VLLM_ASCEND_ENABLE_MLAPO |
enable_mlapo |
"1" → true, "0" → false |
VLLM_ASCEND_ENABLE_NZ |
weight_nz_mode |
Integer (unchanged, field name changed) |
VLLM_ASCEND_ENABLE_FUSED_MC2 |
enable_fused_mc2 |
Integer (unchanged) |
VLLM_ASCEND_FUSION_OP_TRANSPOSE_KV_CACHE_BY_BLOCK |
enable_transpose_kv_cache_by_block |
"1" → true, "0" → false |
VLLM_ASCEND_ENABLE_FLASHCOMM1 |
enable_flashcomm1 |
"1" → true, "0" → false |
How to use¶
With either online mode or offline mode, users can use additional configuration. Take Qwen3 as an example:
Online mode:
Offline mode:
Configuration options¶
The following table lists additional configuration options available in vLLM Ascend:
| Name | Type | Default | Description |
|---|---|---|---|
xlite_graph_config |
dict | {} |
Configuration options for Xlite graph mode |
finegrained_tp_config |
dict | {} |
Configuration options for module tensor parallelism |
ascend_compilation_config |
dict | {} |
Configuration options for ascend compilation |
ascend_warmup_config |
dict | {} |
Configuration options for startup warmup that overlaps weight loading |
eplb_config |
dict | {} |
Runner-specific EPLB extensions. See Expert Parallelism Load Balancer. |
scheduler_config |
dict | {} |
Configuration options for Ascend scheduler extensions, including balance scheduling, recompute scheduling, DyntraLB, ShortRequestFirst, and dynamic chunked pipeline parallel. |
refresh |
bool | false |
Whether to refresh global Ascend configuration content. This is usually used by rlhf or ut/e2e test case. |
dump_config |
dict | None |
Inline msprobe dump configuration. vLLM-Ascend will materialize it to a temporary JSON file and pass that file to the debugger. |
dump_config_path |
str | None |
Configuration file path for msprobe dump (compatible legacy option). |
enable_shared_expert_dp |
bool | False |
Replicate shared-expert weights across TP ranks and run the shared expert with data parallelism. This option is independent of upstream MoE sequence parallelism; either feature or both can be enabled. It improves performance but consumes more memory. |
multistream_overlap_shared_expert |
bool | False |
Whether to enable multi-stream shared expert. This option only takes effect on MoE models with shared experts. |
enable_cpu_binding |
bool | True |
Enables Ascend-native CPU binding on ARM servers. Set to False to disable. See CPU Binding. |
pa_shape_list |
list | [] |
The custom shape list of page attention ops. |
enable_kv_nz |
bool | False |
Whether to enable KV cache NZ layout. This option only takes effect on models using MLA (e.g., DeepSeek). |
c8_enable_reshape_optim |
bool | True |
Whether to use the StoreKVBlock operator to accelerate LightningIndexer C8 cache writes. When enabled, the optimization takes effect only when SFA and LightningIndexer C8 are active on a PD prefill (P) node. |
mc2_comm_alg |
str | "" |
set dispatch/combine op's comm_alg param, only supports ""/"fullmesh"/"hierarchy"/"fullmesh_v2". "hierarchy" is only supported by A2/A3, and "fullmesh_v2" is only supported by A3 now. |
enable_mc2_hierarchy_comm |
bool | False |
Enable dispatch/combine op inter-node communication by ROCE. This param will be deprecated and be replaced by mc2_comm_alg = "hierarchy" |
enable_prefill_mc2 |
bool | False |
Whether to reserve mc2_token_capacity for prefill batches. When enabled, max_num_batched_tokens is used to calculate the mc2_token_capacity instead of the decode-only capacity. In this scenario, the recommended maximum value of max_num_batched_tokens is tp_size * 512. This is a temporary switch; once MC2 operators are complete for all scenarios, this switch will be removed and MC2 will be enabled by default. |
mega_moe_max_tokens |
int | 65536 |
Reference per-rank token capacity after dispatch in the fused MC2/MegaMoe path. It is passed as dispatch_ffn_combine's max_output_size and CANN MegaMoe buffer's max_recv_token_num. If a rank's actual MoE load exceeds this value, precision degradation may occur. The absolute safe upper bound is num_max_tokens_per_rank * int(self.token_dispatcher.ep_world_size) * min(num_topk, expert_per_rank), but using it directly can consume very large device memory. Tune this value based on actual expert load distribution. |
msmonitor_use_daemon |
bool | False |
Whether to use daemon mode for msmonitor. The legacy MSMONITOR_USE_DAEMON environment variable is no longer supported. |
enable_mlapo |
bool | True |
Whether to enable MLAPO (Model Layer-wise Adaptive Parallel Optimization). The legacy VLLM_ASCEND_ENABLE_MLAPO environment variable is no longer supported. |
mlapo_keep_prefill_weights |
bool | False |
When True, keep MLAPO prefill weights on NPU instead of freeing them on kv_consumer (decode-only D) nodes. D nodes have normal local-prefill paths (recompute / fallback / preempt) that crash when the weights are freed (issue #11882). Enable this to trade NPU memory for stability. |
weight_nz_mode |
int | 1 |
Weight NZ mode. 0 disables NZ, 1 enables NZ only for quantized weights, and 2 also enables NZ for BF16/FP16 weights when supported. The legacy VLLM_ASCEND_ENABLE_NZ environment variable is no longer supported. |
enable_fused_mc2 |
int | 0 |
Fused MC2 configuration. 0 disables the fused path, 1 selects dispatch-FFN-combine, and 2 selects CANN MegaMoe when the model and parallel configuration support it. On A5, SiTU models require a CANN MegaMoe wrapper exposing activation and activation_params; both SiTU parameters are bound from the MoE configuration during initialization. The legacy VLLM_ASCEND_ENABLE_FUSED_MC2 environment variable is no longer supported. |
enable_transpose_kv_cache_by_block |
bool | True |
Whether to enable transpose KV cache by block. The legacy VLLM_ASCEND_FUSION_OP_TRANSPOSE_KV_CACHE_BY_BLOCK environment variable is no longer supported. |
enable_dsa_cp |
bool | False |
Whether to enable dsa_cp for DeepSeek V3.2, DeepSeek V4, and other models with the same architecture. This feature requires sequence parallelism to be enabled. Enabling it automatically enables FlashComm. |
enable_flashcomm1 |
bool | False |
Whether to enable SP MoE. The legacy VLLM_ASCEND_ENABLE_FLASHCOMM1 environment variable is kept for compatibility. See Sequence Parallelism. |
enable_pcp_o_proj_weight_sharding |
bool | False |
Whether SFA-PCP shards the O-proj weight across the PCP group at load time and switches between PCP-local and gathered weight views at runtime. This option does not affect DSA-CP, whose original policy remains fixed: prefill gathers the full O-proj weight and decode uses the local weight. This option must be set when the server starts. |
enable_pcp_embedding_lmhead_weight_sharding |
bool | True |
Whether PCP shards embedding and LM Head weights across PCP ranks inside each TP shard. This option is enabled by default, takes effect when PCP size is greater than 1, and is incompatible with fine-grained TP for these modules. |
rejection_sampler_config |
dict | {} |
Configuration options for rejection sampler (block verify and entropy verify). |
dynamic_spec_config |
dict | {} |
Configuration options for Dynamic Speculative Decoding. See Dynamic Speculative Decoding. |
multistream_dsv4_dsa_overlap |
bool | True |
Whether to enable dsa multi-stream overlap for DeepSeek V4. |
rl_config |
dict | {} |
One-click RL mode configuration. See rl_config for all fields, the two deployment modes, usage examples, and the migration guide. |
combine_quant_mode |
int | 0 |
Fused MC2 configuration. This configuration will be passed as the comm_quant_mode argument for the torch_npu.npu_moe_distribute_combine_v2 operator. Please refer to the operator documentation for the valid value range. |
The details of each configuration option are as follows:
[!WARNING] With HDK 0.26.0 or earlier,
c8_enable_reshape_optimmay conflict with pooling models that use AICPU operators. Setc8_enable_reshape_optimtofalseto disable the optimization and avoid the conflict. See issue #15896 for details.
xlite_graph_config
| Name | Type | Default | Description |
|---|---|---|---|
enabled |
bool | False |
Whether to enable Xlite graph mode. See Using XliteGraph for the supported models, the decode-only vs. full-mode distinction, and examples. |
full_mode |
bool | False |
Token budget sizing for the xlite runtime. Batches are routed by token count: those within the budget run on the xlite runtime, larger ones fall back to the runnable under ACLGraph. By default (False), the budget is sized for decode steps (max_num_seqs × (1 + num_speculative_tokens)), so prefill and large mixed batches typically fall back. When True, the budget is max_num_batched_tokens, so xlite handles prefill and decode alike, ACLGraph capture is not used, and --enforce-eager is recommended (unless speculative decoding is configured, etc.). Since v0.28.0, routing is by token count instead of the batch attention state. |
finegrained_tp_config
| Name | Type | Default | Description |
|---|---|---|---|
lmhead_tensor_parallel_size |
int | 0 |
The custom tensor parallel size of lm_head. |
oproj_tensor_parallel_size |
int | 0 |
The custom tensor parallel size of o_proj. |
embedding_tensor_parallel_size |
int | 0 |
The custom tensor parallel size of embedding. |
mlp_tensor_parallel_size |
int | 0 |
The custom tensor parallel size of mlp. |
ascend_compilation_config
| Name | Type | Default | Description |
|---|---|---|---|
enable_npugraph_ex |
bool | True |
Whether to enable npugraph_ex backend. |
enable_static_kernel |
bool | False |
Whether to enable static kernel. Suitable for scenarios where shape changes are minimal and some time is available for static kernel compilation. |
enable_super_kernel |
bool | Inherits enable_static_kernel |
Whether to enable Super Kernel optimization. When omitted, it follows enable_static_kernel; set it explicitly to override that behavior. Super Kernel requires static kernel. |
fuse_norm_quant |
bool | True |
Whether to enable fuse_norm_quant pass. |
fuse_qknorm_rope |
bool | True |
Whether to enable fuse_qknorm_rope pass. If Triton is not in the environment, set it to False. |
fuse_muls_add |
bool | True |
Whether to enable fuse_muls_add pass. |
ascend_warmup_config
Both warmups run on a background thread during weight loading. The worker waits for them at the end of model loading, before memory profiling and KV cache allocation.
| Name | Type | Default | Description |
|---|---|---|---|
enable_early_kernel_warmup |
bool | False |
Compile the rejection sampler, penalty, and RMS norm Triton warmup kernels while weights load, so the regular kernel warmup hits the Triton cache. |
enable_early_nz_warmup |
bool | False |
Pay the one-time lazy initialization of the first NZ format cast while weights load. Independent of the quantization scheme. |
eplb_config
The accepted fields depend on the model runner:
- Model Runner V2 accepts
load_collection_phaseandstair_confighere. Configure upstream EPLB through--enable-eplband--eplb-config. Ascend uses the STAIR policy by default and asynchronous Gloo movement. Selectdefaultorstaironly through the upstream--eplb-config.policyoption. See the EPLB user guide for the advanced STAIR fields and their defaults. - Model Runner V1 accepts the legacy fields below except
load_collection_phase. MRv1 does not accept upstream--enable-eplbon Ascend.
Mixing the two schemas fails during startup instead of silently ignoring configuration.
| Name | Type | Default | Description |
|---|---|---|---|
dynamic_eplb |
bool | False |
MRv1 only. Whether to enable legacy dynamic EPLB. |
expert_map_path |
str | None |
MRv1 only. Load a recorded static expert map. |
expert_heat_collection_interval |
int | 600 |
MRv1 only. Number of forward iterations used to collect expert heat. |
algorithm_execution_interval |
int | 50 |
MRv1 only. Interval allowed for the EPLB worker to finish its CPU task. |
expert_map_record_path |
str | None |
MRv1 only. Save the calculated expert map to the specified JSON path. |
num_redundant_experts |
int | 0 |
MRv1 only in this table. Configure the MRv2 value through upstream --eplb-config. |
eplb_policy_type |
int | 2 |
MRv1 only. EPLB policy: 0=Random, 1=DefaultEplb, 2=SwiftBalanceEplb, 3=FlashLB. |
eplb_heat_collection_stage |
str | "all" |
MRv1 only. Select "all", "prefill", or "decode" heat collection. |
load_collection_phase |
str | "all" |
MRv2 only. Select "all", "prefill", or "decode" load submission. Any batch containing a prefill request is classified entirely as prefill. |
scheduler_config
The legacy top-level enable_balance_scheduling, recompute_scheduler_enable, short_request_first_config, and profiling_chunk_config keys remain supported during the migration period, but are deprecated. If both formats provide the same field, the value in scheduler_config takes precedence.
| Name | Type | Default | Description |
|---|---|---|---|
enable_balance_scheduling |
bool | False |
Whether to enable balance scheduling. The legacy VLLM_ASCEND_BALANCE_SCHEDULING environment variable is no longer supported. |
recompute_scheduler_enable |
bool | False |
Whether to enable the recompute scheduler. Only valid on PD-disaggregated D nodes (kv_role is kv_consumer). Do not enable on P nodes or in PD-mixed mode (no kv_transfer_config, kv_role is kv_producer, or kv_role is kv_both); startup will fail with a clear error. |
profiling_chunk_config |
dict | {} |
Configuration options for dynamic chunked pipeline parallel. See Dynamic Chunked Pipeline Parallel for details. |
short_request_first_config |
dict | {} |
Configuration options for ShortRequestFirst prefill scheduling on FCFS synchronous or asynchronous, PD-prefill (P), or PD-mixed nodes. |
batch_job_sched_config |
dict | {} |
Configuration options for the batch-job-aware scheduler. See Batch-Job-Aware Scheduler for details. |
dyntra_lb_config |
dict | {} |
Configuration options for DyntraLB load balancing on PD-disaggregated decode nodes. |
scheduler_config.profiling_chunk_config
| Name | Type | Default | Description |
|---|---|---|---|
enabled |
bool | False |
Whether to enable dynamic chunked pipeline parallel. Requires pipeline-parallel-size > 1. |
smooth_factor |
float | 1.0 |
Smoothing factor (0 < x ≤ 1.0). Higher values trust the dynamic prediction more; 0.0 disables dynamic adjustment. |
min_chunk |
int | 4096 |
Minimum chunk size for dynamic calculation. Should be smaller than max-num-batched-tokens. |
need_timing |
bool | True | Enable/disable Online Calibration |
max_fit_chunk |
int | 30 | Number of chunk-time data for Online Calibration |
scheduler_config.dyntra_lb_config
DyntraLB balances decode requests across data-parallel ranks. It is supported only on
PD-disaggregated decode nodes (kv_role="kv_consumer") with data_parallel_size > 1.
dyntra_lb_config.enabled and recompute_scheduler_enable are independent sibling
settings; enabling both selects the combined DyntraLB recompute scheduler.
| Name | Type | Default | Description |
|---|---|---|---|
enabled |
bool | False |
Enable DyntraLB and select a DyntraLB-aware scheduler. |
mode |
str | "dynamic" |
Use "static" or "dynamic" activation. |
start_step |
int | 250 |
First completed engine-step snapshot allowed to generate a plan. |
end_step |
int | -1 |
Exclusive final snapshot step; -1 means no upper bound. |
bubble_threshold |
float | 5.0 |
Minimum maximum-to-average rank-load difference required to modify scheduling. Values greater than or equal to 1 are KV-cache blocks; values below 1 are normalized ratios. |
long_req_block_threshold |
int | 700 |
In dynamic mode, a newly added request above this block count activates balancing. The default threshold corresponds to approximately 89,600 tokens when block_size=128. |
dynamic_max_step |
int | 256 |
Stop dynamic balancing after this many active steps without another newly added long request. |
enable_diagnostics |
bool | False |
Enable verbose logs for feature validation and debugging only. It is disabled by default and should remain disabled in production. |
rejection_sampler_config
Note: Both block verify and entropy verify improve speculative decoding performance (higher acceptance rate, lower latency) at the cost of reduced sampling precision. A larger
posterior_alphamakes the adjustment more aggressive — it further lowers the acceptance threshold for high-entropy tokens, improving throughput but degrading output quality. Users should tune these parameters based on their specific model weights and application scenario to find the right trade-off between performance and precision.
| Name | Type | Default | Description |
|---|---|---|---|
enable_block_verify |
bool | False |
Whether to enable block verify mode. Block verify evaluates all draft tokens as a block using cumulative probability products, which can improve acceptance rate. |
enable_entropy_verify |
bool | False |
Whether to enable entropy verify mode. Entropy verify adjusts the acceptance threshold based on the entropy of the target distribution — higher entropy (uncertain) tokens get a lower threshold (easier to accept), while lower entropy (confident) tokens get a stricter threshold. |
posterior_threshold |
float | 0.95 |
Upper bound for the entropy-adjusted acceptance threshold. Must be in (0, 1]. The effective threshold is min(exp(-entropy * posterior_alpha), posterior_threshold). |
posterior_alpha |
float | 0.4 |
Scaling factor for entropy in the threshold computation. Must be >= 0. Higher values make the threshold more sensitive to entropy — high-entropy tokens become much easier to accept, improving performance but reducing precision. |
dynamic_spec_config
Note: This is an exploratory feature for model runner v1. Supported methods are
"dspark"(DSpark confidence head) and"dflash"(head-free; uses max-softmax over draft logits as a confidence proxy). You still need a matchingspeculative_config(method: "dspark"or"dflash");dynamic_spec_configonly controls how many drafted tokens are verified per request. See Dynamic Speculative Decoding for usage and limitations.
| Name | Type | Default | Description |
|---|---|---|---|
method |
str | None |
Dynamic method name. Supported values: "dspark", "dflash". Omit or set to None to disable. |
method_params |
dict | {} |
Method-specific hyperparameters. When empty, each method falls back to its built-in defaults. |
dynamic_spec_config.method_params (when method is "dspark" or "dflash")
dspark and dflash share the same scheduling hyperparameters. The difference is only how per-token acceptance confidence is estimated: DSpark uses its confidence head (sigmoid), while DFlash (head-free) uses max(softmax(logits)) of the drafted token.
| Name | Type | Default | Description |
|---|---|---|---|
initial_verify_budget_per_req |
int | 5 |
Initial per-request verify budget before the first recompute. |
budget_update_interval |
int | 16 |
Recompute the shared verify budget every N decode steps. |
budget_threshold |
float | 0.3 |
Cumulative survival-probability threshold used when estimating the mean verify budget. |
min_verify_tokens |
int | 1 |
Minimum number of draft tokens verified per request. |
scheduler_config.short_request_first_config
ShortRequestFirst is a waiting-queue policy for FCFS synchronous or asynchronous scheduling on prefill and PD-mixed paths. It does not support batch-job-aware, profiling-chunk, or PD-disaggregated D-node scheduling. See ShortRequestFirst Prefill Scheduling for usage, behavior, and tuning guidance.
| Name | Type | Default | Description |
|---|---|---|---|
enabled |
bool | False |
Whether to enable ShortRequestFirst scheduling. |
threshold |
int | 256 |
Prompt-length threshold (tokens). Requests with num_prompt_tokens <= threshold are treated as short prefills and prioritized over long prefills. |
long_max_wait_ms |
float | 0.0 |
Maximum time a long prefill may wait behind short prefills before it can be promoted ahead of them. 0 disables long-request promotion and keeps strict short-request priority. |
scheduler_config.batch_job_sched_config
| Name | Type | Default | Description |
|---|---|---|---|
enabled |
bool | false |
Enable the batch-job-aware scheduler. |
max_jobs |
int | 20 |
Maximum number of tracked jobs. 0 means unlimited. |
reserve_margin_blocks |
int | 2 |
Extra block margin added to the KV cache reserve as safety buffer. |
reserve_max_blocks |
int | 8 |
Maximum number of blocks that can be reserved. |
low_available_tokens_threshold |
int | 4096 |
Threshold for prioritising long vs short decode jobs. When available tokens > threshold, long decode jobs are prioritised; when ≤ threshold, short decode jobs are prioritised. |
short_decode_token_threshold |
int | 32 |
Threshold for classifying a job as "short decode". |
rl_config
rl_config is a one-click RL mode switch. When enabled is true, it refreshes the global Ascend configuration on every initialization, forces AscendConfig.weight_nz_mode=0, sets VLLM_SERVER_DEV_MODE=1, and removes the expandable_segments entry from PYTORCH_NPU_ALLOC_CONF with an informational log. These fixed RL behaviors are not configurable as rl_config sub-fields. When enabled is false, all other sub-fields are ignored.
When RL mode is enabled, its fixed NZ setting takes precedence over the top-level weight_nz_mode configuration. VLLM_BATCH_INVARIANT=1 remains enabled when rl_config.enable_batch_invariant is false.
| Name | Type | Default | Description |
|---|---|---|---|
enabled |
bool | false |
Master switch for RL mode. When true, all RL best-practice defaults below are applied. |
sleep_mode_extra_cleanup |
bool | false |
Same-device mode. Enables HCCL process-group release + ACL graph workspace cleanup during sleep, returning more NPU memory to the trainer at the cost of increased wakeup latency. This option is available only under rl_config; the former top-level enable_sleep_mode_extra_cleanup key has been removed. |
enable_training_consistency |
bool | false |
Both modes. Enables the FA3 attention backend used for training-inference consistency. Requires the flash_attn_npu_v3 package and does not implicitly enable batch invariance. |
enable_batch_invariant |
bool | false |
Both modes. Provides an additional way to enable batch-invariant deterministic computation: when true, sets VLLM_BATCH_INVARIANT=1 before workers are started. Worker-side batch-invariant initialization then sets HCCL_DETERMINISTIC=strict and LCCL_DETERMINISTIC=1. When false, an existing VLLM_BATCH_INVARIANT environment setting is preserved. Requires building vllm-ascend from source with COMPILE_CUSTOM_KERNELS=1; RL mode already forces weight_nz_mode=0. |
Deployment modes
- Same-device mode (sleep/wake + IPC weight transfer): the inference engine and trainer share the same NPU card. Use
enable_sleep_mode=Trueand optionallysleep_mode_extra_cleanup=Truefor extra memory. - Cross-device mode (pause/resume + HCCL weight transfer): the inference engine and trainer use different NPU cards. Use
--weight-transfer-config '{"backend": "nccl"}'.
Example (online, same-device):
vllm serve DeepSeek-V4 \
--enable-sleep-mode \
--enable-return-routed-experts \
--weight-transfer-config '{"backend": "ipc"}' \
--additional-config '{"rl_config": {"enabled": true, "enable_training_consistency": true, "enable_batch_invariant": true, "sleep_mode_extra_cleanup": true}}'
Example (online, cross-device):
vllm serve DeepSeek-V4 \
--weight-transfer-config '{"backend": "nccl"}' \
--additional-config '{"rl_config": {"enabled": true}}'
Example (offline):
from vllm import LLM
llm = LLM(
model="DeepSeek-V4",
enable_sleep_mode=True, # same-device mode only
additional_config={
"rl_config": {
"enabled": True,
"enable_training_consistency": True,
"enable_batch_invariant": True,
},
},
)
Migration guide
| Before | After |
|---|---|
export VLLM_ASCEND_ENABLE_NZ=0 |
"rl_config": {"enabled": true} |
export VLLM_SERVER_DEV_MODE=1 |
"rl_config": {"enabled": true} |
export VLLM_BATCH_INVARIANT=1 |
"rl_config": {"enabled": true, "enable_batch_invariant": true} |
--attention-backend FLASH_ATTN for FA3 consistency mode |
"rl_config": {"enabled": true, "enable_training_consistency": true} |
top-level "weight_nz_mode": 0 for RL |
"rl_config": {"enabled": true} |
top-level "enable_sleep_mode_extra_cleanup": true |
"rl_config": {"enabled": true, "sleep_mode_extra_cleanup": true} |
Example¶
An example of additional configuration is as follows:
{
"finegrained_tp_config": {
"lmhead_tensor_parallel_size": 8,
"oproj_tensor_parallel_size": 8,
"embedding_tensor_parallel_size": 8,
"mlp_tensor_parallel_size": 8,
},
"enable_kv_nz": False,
"multistream_overlap_shared_expert": True,
"rejection_sampler_config": {
"enable_block_verify": True,
"enable_entropy_verify": True,
"posterior_threshold": 0.95,
"posterior_alpha": 0.4,
},
"dynamic_spec_config": {
"method": "dspark",
"method_params": {
"initial_verify_budget_per_req": 5,
"budget_update_interval": 50,
"budget_threshold": 0.7,
},
},
"refresh": False
}
KV pipeline parallelism (KVPP)¶
Set enable_kvpp: true in --additional-config to distribute persistent MLA
KV-cache layers across TP and (with Model Runner V2) PCP ranks within the same
DP replica and PP stage. Group size is TP x PCP: TP4 + PCP2 uses eight ranks;
TP1 + PCP2 also enables KVPP. DP replicas and PP stages use separate groups.
PCP requires VLLM_USE_V2_MODEL_RUNNER=1. KVPP still requires eager execution and
non-hybrid MLA, and does not support DCP or KV transfer connectors. PCP gathers
prefill KV before cache writes; layer broadcasts restore prior-forward cache
contents before attention. The broadcast decision uses the global scheduled
batch, not PCP-local segment offsets.