跳转至

GLM-5 与 GLM-5.1

1 引言

本文档同时适用于 GLM-5 和 GLM-5.1。除非另有说明,本文档中所有关于 GLM-5 的描述、配置和部署流程同样适用于 GLM-5.1。为简洁起见,下文统一使用 GLM-5 指代 GLM-5 和 GLM-5.1。

GLM-5 采用混合专家(MoE)架构,面向复杂的系统工程和长周期智能体任务。

GLM-5 模型首次在 vllm-ascend:v0.17.0rc1 中得到支持(对于 Ascend 950DT系列产品,该模型从 vllm-ascend:v0.23.0rc1 开始支持),所有 v0.17.0rc1 及之后的版本 均可稳定运行。如需使用最新特性(例如 PD 分离、MTP),建议使用最新的候选版本或正式版本。transformers 的版本需要升级到 5.2.0 或更高版本。

本文档将展示该模型的主要验证步骤,包括支持的特性、特性配置、环境准备、单节点和多节点部署、精度及性能评估。

2 支持的特性

请参阅支持的功能以获取该模型支持的功能矩阵。

请参阅功能指南以获取该功能的配置。

3 前提条件

3.1 模型权重

权重版本 硬件要求 下载链接
GLM-5-w4a8(量化版本) 1 个 Atlas 800 A3(128GB × 8)节点或 2 个 Atlas 800 A2(64GB × 8)节点 ModelScope
GLM-5-w8a8(量化版本) 1 个 Atlas 800 A3(128GB × 8)节点或 2 个 Atlas 800 A2(64GB × 8)节点 ModelScope
GLM-5.1-w4a8(量化版本) 1 个 Atlas 800 A3(128GB × 8)节点或 2 个 Atlas 800 A2(64GB × 8)节点 Modelers
GLM-5.1-w8a8(量化版本) 1 个 Atlas 800 A3(128GB × 8)节点或 2 个 Atlas 800 A2(64GB × 8)节点 Modelers
GLM-5.1-w8a8c8(量化版本) 该权重已在 Atlas 800 A3 上验证,推荐使用。 Modelers
GLM-5.1-w4a4(Ascend 950DT系列产品 mxfp4 量化) 该权重已在 Ascend 950DT系列产品上验证,推荐使用。 ModelScope

建议将模型权重下载到多节点的共享目录,例如 /root/.cache/。

路径说明:将模型权重下载到您选择的目录并记录该路径。确保后续部署命令中的模型路径与该目录一致。

3.2 验证多节点通信(可选)

如果需要多节点部署,请按照验证多节点通信环境指南进行通信验证。

4 安装

4.1 Docker 镜像安装

您可以直接使用我们的官方 Docker 镜像来运行 GLM-5/5.1。

在每个节点上启动docker镜像。

export IMAGE=quay.io/ascend/vllm-ascend:v0.23.0-a5
export NAME=vllm-ascend

docker run --rm \
--name $NAME \
--net=host \
--shm-size=1g \
--device /dev/davinci0 \
--device /dev/davinci1 \
--device /dev/davinci2 \
--device /dev/davinci3 \
--device /dev/davinci4 \
--device /dev/davinci5 \
--device /dev/davinci6 \
--device /dev/davinci7 \
--device /dev/davinci_manager \
--device /dev/hisi_hdc \
--device /dev/ummu \
--device /dev/uburma \
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
-v /etc/ascend_install.info:/etc/ascend_install.info \
-v /etc/hccl_rootinfo.json:/etc/hccl_rootinfo.json \
-v /etc/hixlep/:/etc/hixlep/ \
-v /root/.cache:/root/.cache \
-v /usr/local/sbin:/usr/local/sbin \
-v /usr/local/dcmi:/usr/local/dcmi \
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
-v /usr/local/sbin/npu-smi:/usr/local/sbin/npu-smi \
-v /usr/bin/urma_admin:/usr/bin/urma_admin \
-v /lib/route.conf:/lib/route.conf \
-v /usr/lib64:/usr/lib64 \
-itd $IMAGE bash

在每个节点上启动docker镜像。

export IMAGE=quay.io/ascend/vllm-ascend:v0.23.0-a3
export NAME=vllm-ascend

# Run the container using the defined variables
# Note: If you are running bridge network with docker, please expose available ports for multiple nodes communication in advance
docker run --rm \
--name $NAME \
--net=host \
--shm-size=1g \
--device /dev/davinci0 \
--device /dev/davinci1 \
--device /dev/davinci2 \
--device /dev/davinci3 \
--device /dev/davinci4 \
--device /dev/davinci5 \
--device /dev/davinci6 \
--device /dev/davinci7 \
--device /dev/davinci8 \
--device /dev/davinci9 \
--device /dev/davinci10 \
--device /dev/davinci11 \
--device /dev/davinci12 \
--device /dev/davinci13 \
--device /dev/davinci14 \
--device /dev/davinci15 \
--device /dev/davinci_manager \
--device /dev/devmm_svm \
--device /dev/hisi_hdc \
-v /usr/local/dcmi:/usr/local/dcmi \
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
-v /etc/ascend_install.info:/etc/ascend_install.info \
-v /root/.cache:/root/.cache \
-it $IMAGE bash

在每个节点上启动docker镜像。

export IMAGE=quay.io/ascend/vllm-ascend:v0.23.0
docker run --rm \
    --name vllm-ascend \
    --shm-size=1g \
    --net=host \
    --device /dev/davinci0 \
    --device /dev/davinci1 \
    --device /dev/davinci2 \
    --device /dev/davinci3 \
    --device /dev/davinci4 \
    --device /dev/davinci5 \
    --device /dev/davinci6 \
    --device /dev/davinci7 \
    --device /dev/davinci_manager \
    --device /dev/devmm_svm \
    --device /dev/hisi_hdc \
    -v /usr/local/dcmi:/usr/local/dcmi \
    -v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
    -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
    -v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
    -v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
    -v /etc/ascend_install.info:/etc/ascend_install.info \
    -v /root/.cache:/root/.cache \
    -it $IMAGE bash

如果要部署多节点环境,需要在每个节点上进行环境设置。

要验证环境是否安装成功,请参阅安装。

4.2 源码安装

此外,如果您不想使用上述 Docker 镜像,也可以从源码构建所有内容:

  • 从源码安装vllm-ascend,请参阅安装。

如果要部署多节点环境,需要在每个节点上进行环境设置。

5 在线服务部署

5.1 单节点在线部署

  • 量化模型 glm-5-w4a4 可部署在 1 个 Ascend 950DT系列产品(96GB × 8)上。

运行以下脚本执行在线推理。

常见问题提示:如果遇到问题,请参阅常见问题。

#!/usr/bin/env bash
source /root/.bashrc
# this obtained through ifconfig
# nic_name is the network interface name corresponding to local_ip of the current node
nic_name="xxx"
local_ip="xxx"

export HCCL_BUFFSIZE=400
export HCCL_IF_IP=$local_ip
export HCCL_INTRA_ROCE_ENABLE=0
export HCCL_OP_EXPANSION_MODE="AIV"
export HCCL_SOCKET_IFNAME=$nic_name
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
export DYNAMIC_EPLB="true"
export GLOO_SOCKET_IFNAME=$nic_name
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
export PROMETHEUS_MULTIPROC_DIR=/dev/shm/vllm_metrics && mkdir -p $PROMETHEUS_MULTIPROC_DIR
export HCCL_DFS_CONFIG="task_exception:off,inconsistent_check:off"
export VLLM_ASCEND_ENABLE_PREFETCH_MLP=1

# Ensure the model path matches the directory recorded during download
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM5-w4a4 \
--host 0.0.0.0 \
--port 8000 \
--data-parallel-size 1 \
--tensor-parallel-size 8 \
--seed 1024 \
--served-model-name glm-5 \
--enable-expert-parallel \
--max-num-seqs 128 \
--max-model-len 202752 \
--max-num-batched-tokens 8192 \
--trust-remote-code \
--enable-prefix-caching \
--gpu-memory-utilization 0.95 \
--quantization ascend \
--enable-auto-tool-choice \
--tool-call-parser glm47 \
--reasoning-parser glm45 \
--kv-cache-dtype fp8 \
--attention_config.indexer_kv_dtype fp8 \
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
--hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
--speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp", "enforce_eager": true}' \
--additional-config '{"enable_cpu_binding": "True", "multistream_overlap_shared_expert": "True", "enable_dsa_cp": true, "eplb_config": {"dynamic_eplb": true, "expert_heat_collection_interval": 50, "algorithm_execution_interval": 5, "eplb_policy_type": 2, "num_redundant_experts": 0}, "enable_flashcomm1": true}'
  • 量化模型glm-5-w4a8和glm-5.1-w4a8可以部署在1台Atlas 800 A3(128GB × 8)上。

运行以下脚本执行在线推理。

常见问题提示:如果遇到问题,请参阅常见问题。

# The version of transformers needs to be upgraded to 5.2.0.
# pip install transformers==5.2.0 --upgrade

export HCCL_BUFFSIZE=200
export HCCL_OP_EXPANSION_MODE="AIV"
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True

# Ensure the model path matches the directory recorded during download
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM5-w4a8 \
--host 0.0.0.0 \
--port 8000 \
--data-parallel-size 1 \
--tensor-parallel-size 16 \
--enable-expert-parallel \
--seed 1024 \
--served-model-name glm-5 \
--max-num-seqs 16 \
--max-model-len 200000 \
--max-num-batched-tokens 4096 \
--trust-remote-code \
--gpu-memory-utilization 0.95 \
--quantization ascend \
--enable-chunked-prefill \
--enable-prefix-caching \
--additional-config '{"multistream_overlap_shared_expert": true, "enable_balance_scheduling": true, "enable_flashcomm1": true}' \
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
--hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
--speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp", "enforce_eager": true}'
  • 量化模型 glm-5-w4a8 可部署在1台Atlas 800 A2(64GB × 8)上。

运行以下脚本执行在线推理。

常见问题提示:如果遇到问题,请参阅常见问题。

export HCCL_BUFFSIZE=200
export HCCL_OP_EXPANSION_MODE="AIV"
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True

# Ensure the model path matches the directory recorded during download
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM5-w4a8 \
--host 0.0.0.0 \
--port 8000 \
--data-parallel-size 1 \
--tensor-parallel-size 8 \
--enable-expert-parallel \
--seed 1024 \
--served-model-name glm-5 \
--max-num-seqs 8 \
--max-model-len 32768 \
--max-num-batched-tokens 4096 \
--trust-remote-code \
--gpu-memory-utilization 0.95 \
--quantization ascend \
--enable-chunked-prefill \
--enable-prefix-caching \
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
--additional-config '{"multistream_overlap_shared_expert": true, "enable_balance_scheduling": true, "enable_flashcomm1": true}' \
--hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
--speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp", "enforce_eager": true}'

关键参数说明:

以下仅描述该模型/场景特有的关键参数。max-model-len 和 max-num-seqs 需要根据实际使用场景进行设置。

模型特有参数:

  • --enable-expert-parallel:对于GLM-5的MoE架构必须启用。
  • --tensor-parallel-size 16 / --tensor-parallel-size 8:每个DP rank内的张量并行。对于A3(16个NPU),使用tp16;对于A2(8个NPU),使用tp8。
  • --quantization ascend:为w4a8/w8a8量化权重启用Ascend量化。
  • --speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp", "enforce_eager": true}':使用GLM-5的DeepSeek风格MTP草稿模型启用多令牌预测(MTP)投机解码。num_speculative_tokens(3-5)控制每步推测的令牌数量;enforce_eager: true是必需的,因为GLM-5不支持图模式投机解码。
  • --enable-chunked-prefill / --enable-prefix-caching:推荐用于长上下文和多用户场景——chunked prefill将长提示拆分以改善TTFT,prefix caching为共享前缀(如系统提示)重用KV缓存。
  • --compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}':仅对解码阶段启用图捕获,通过减少内核启动开销来提升解码性能。
  • --additional-config '{"multistream_overlap_shared_expert": true}':在额外的流上重叠共享专家计算。注意:当"enable_fused_mc2": true时自动禁用,因为这两种优化存在冲突。

关键的--additional-config字段:

  • "enable_flashcomm1": true:启用FlashComm优化以减少通信开销(主要有利于prefill路径)。启用FlashComm后,layer_sharding不能包含o_proj。
  • "enable_mlapo": true:启用MLA预处理融合算子(MlaPreprocessOperation)。对于w8a8模型默认启用——显著提升Decode性能但消耗更多NPU内存;如果内存优先,请设置"enable_mlapo": false。推荐用于w8a8;w4a8可能无法受益。
  • "enable_balance_scheduling": true:启用平衡调度以提升v1调度器中的输出吞吐量并降低TPOT。

单节点性能调优说明:

  • 对于低延迟场景,使用 dp1tp16(data-parallel-size 1,tensor-parallel-size 16),并考虑减小 --max-num-seqs 和 --max-num-batched-tokens。
  • 对于高吞吐场景,增大 --max-num-seqs 并启用 --enable-prefix-caching。
  • 对于长上下文场景(例如 200k),使用 w4a8 权重(为 KV cache 留出更多内存),并将 --max-model-len 设置为所需的上下文长度。考虑启用 --enable-chunked-prefill。
  • 如果遇到 OOM,请减小 --gpu-memory-utilization、--max-num-seqs 或 --max-model-len。禁用 "enable_mlapo" 也可以减少内存占用(但会牺牲性能)。

5.2 多节点部署

如果要部署多节点环境,需要按照验证多节点通信环境进行多节点通信验证。

常见问题提示:如果遇到问题,请参阅常见问题。

高吞吐量场景(DP8 TP4)

  • glm-5.1-w8a8c8:可以部署在2台Atlas 800 A3(128GB × 8)上用于高吞吐量场景。

分别在两个节点上运行以下脚本。

节点 0

# this obtained through ifconfig
# nic_name is the network interface name corresponding to local_ip of the current node
nic_name="xxx"
local_ip="xxx"
export HCCL_BUFFSIZE=400
export HCCL_IF_IP=$local_ip
export HCCL_OP_EXPANSION_MODE="AIV"
export HCCL_SOCKET_IFNAME=$nic_name
export GLOO_SOCKET_IFNAME=$nic_name
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
# Ensure the model path matches the directory recorded during download
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM-5.1-W8A8C8-MTP \
--host 0.0.0.0 \
--port 8000 \
--data-parallel-size 8 \
--data-parallel-size-local 4 \
--data-parallel-address $local_ip \
--enable-expert-parallel \
--data-parallel-rpc-port 12980 \
--tensor-parallel-size 4 \
--hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
--seed 1024 \
--served-model-name glm-5 \
--tool-call-parser glm47 \
--reasoning-parser glm45 \
--enable-auto-tool-choice \
--trust-remote-code \
--gpu-memory-utilization 0.92 \
--quantization ascend \
--enable-chunked-prefill \
--enable-prefix-caching \
--async-scheduling \
--kv-cache-dtype int8 \
--attention_config.indexer_kv_dtype int8 \
--additional-config '{"enable_dsa_cp": true, "enable_balance_scheduling": true, "fuse_muls_add": true, "enable_flashcomm1": true, "enable_fused_mc2": true, "enable_mlapo": true}' \
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
--speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp","enforce_eager":true}'

节点 1

# this obtained through ifconfig
# nic_name is the network interface name corresponding to local_ip of the current node
nic_name="xxx"
local_ip="xxx"
# IP of node 0 (the data parallel master node), must be consistent with the local_ip of node 0
node0_ip="xxxx"

export HCCL_BUFFSIZE=400
export HCCL_IF_IP=$local_ip
export HCCL_OP_EXPANSION_MODE="AIV"
export HCCL_SOCKET_IFNAME=$nic_name
export GLOO_SOCKET_IFNAME=$nic_name
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
# Ensure the model path matches the directory recorded during download
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM-5.1-W8A8C8-MTP \
--host 0.0.0.0 \
--port 8000 \
--headless \
--data-parallel-size 8 \
--data-parallel-size-local 4 \
--data-parallel-start-rank 4 \
--data-parallel-address $node0_ip \
--enable-expert-parallel \
--data-parallel-rpc-port 12980 \
--tensor-parallel-size 4 \
--hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
--seed 1024 \
--served-model-name glm-5 \
--tool-call-parser glm47 \
--reasoning-parser glm45 \
--enable-auto-tool-choice \
--trust-remote-code \
--gpu-memory-utilization 0.92 \
--quantization ascend \
--enable-chunked-prefill \
--enable-prefix-caching \
--async-scheduling \
--kv-cache-dtype int8 \
--attention_config.indexer_kv_dtype int8 \
--additional-config '{"enable_dsa_cp": true, "enable_balance_scheduling": true, "fuse_muls_add": true, "enable_flashcomm1": true, "enable_fused_mc2": true, "enable_mlapo": true}' \
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
--speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp","enforce_eager":true}'

注意:

  • 当测试的前缀缓存命中率 > 0 时,保留 --enable-prefix-caching(如上述脚本所示);当命中率为 0 时,将其替换为 --no-enable-prefix-caching。
  • "enable_fused_mc2": true 与 "multistream_overlap_shared_expert": true 冲突——当启用 fused MC2 时,运行时自动禁用 multistream_overlap_shared_expert。

分别在两个节点上运行以下脚本。

节点 0

# this obtained through ifconfig
# nic_name is the network interface name corresponding to local_ip of the current node
nic_name="xxx"
local_ip="xxx"

# The value of node0_ip must be consistent with the value of local_ip set in node0 (master node)
node0_ip="xxx"

export HCCL_BUFFSIZE=200
export HCCL_IF_IP=$local_ip
export HCCL_OP_EXPANSION_MODE="AIV"
export HCCL_SOCKET_IFNAME=$nic_name
export GLOO_SOCKET_IFNAME=$nic_name
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True

# Ensure the model path matches the directory recorded during download
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM5-w4a8 \
--host 0.0.0.0 \
--port 8000 \
--data-parallel-size 2 \
--data-parallel-size-local 1 \
--data-parallel-address $node0_ip \
--data-parallel-rpc-port 13389 \
--tensor-parallel-size 8 \
--quantization ascend \
--seed 1024 \
--served-model-name glm-5 \
--enable-expert-parallel \
--max-num-seqs 2 \
--max-model-len 131072 \
--max-num-batched-tokens 4096 \
--trust-remote-code \
--gpu-memory-utilization 0.95 \
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
--additional-config '{"multistream_overlap_shared_expert": true, "enable_balance_scheduling": true, "enable_flashcomm1": true}' \
--hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
--speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp", "enforce_eager": true}'

节点 1

# this obtained through ifconfig
# nic_name is the network interface name corresponding to local_ip of the current node
nic_name="xxx"
local_ip="xxx"

# The value of node0_ip must be consistent with the value of local_ip set in node0 (master node)
node0_ip="xxx"

export HCCL_BUFFSIZE=200
export HCCL_IF_IP=$local_ip
export HCCL_OP_EXPANSION_MODE="AIV"
export HCCL_SOCKET_IFNAME=$nic_name
export GLOO_SOCKET_IFNAME=$nic_name
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True

# Ensure the model path matches the directory recorded during download
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM5-w4a8 \
--host 0.0.0.0 \
--port 8000 \
--headless \
--data-parallel-size 2 \
--data-parallel-size-local 1 \
--data-parallel-start-rank 1 \
--data-parallel-address $node0_ip \
--data-parallel-rpc-port 13389 \
--tensor-parallel-size 8 \
--quantization ascend \
--seed 1024 \
--served-model-name glm-5 \
--enable-expert-parallel \
--max-num-seqs 2 \
--max-model-len 131072 \
--max-num-batched-tokens 4096 \
--trust-remote-code \
--gpu-memory-utilization 0.95 \
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
--additional-config '{"multistream_overlap_shared_expert": true, "enable_balance_scheduling": true, "enable_flashcomm1": true}' \
--hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
--speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp", "enforce_eager": true}'

多节点部署的关键参数说明:

除了单节点在线部署中描述的所有单节点参数外,以下参数专门用于多节点部署:

网络和数据并行配置:

  • HCCL_IF_IP、GLOO_SOCKET_IFNAME、HCCL_SOCKET_IFNAME:多节点通信的网络接口配置。将 nic_name 设置为网络接口名称(通过 ifconfig 获取),将 local_ip 设置为当前节点的 IP 地址。每个节点上都必须正确配置这些参数,才能实现成功的多节点通信。
  • --data-parallel-size:所有节点上的数据并行 rank 总数。对于 2 节点部署,通常设置为 2。
  • --data-parallel-size-local:当前节点上的数据并行 rank 数量。通常设置为 1(每个节点一个 DP rank)。
  • --data-parallel-address:数据并行主节点(节点 0)的 IP 地址。必须与主节点的 local_ip 匹配。
  • --data-parallel-rpc-port:数据并行主节点通信的 RPC 端口。所有节点上必须相同。
  • --headless:表示这是非主节点。不要在节点 0 上使用。
  • --data-parallel-start-rank:此节点上数据并行 rank 的起始 rank 偏移量。节点 0 使用 0;节点 1 使用节点 0 上的 DP rank 数量(例如,DP2 为 1,DP8 为 4)。

多节点性能调优:

  • 对于低延迟多节点场景,保持 --data-parallel-size-local 1 以最小化跨节点通信。
  • --max-num-seqs 应根据模型加载后的可用 KV 缓存内存进行调整。对于 A3 上的 w8a8c8 198K 高吞吐量场景,建议设置为 6;对于 198K 低延迟场景,建议设置为 16。对于 A2 上长上下文的 w4a8 多节点场景,从 2 开始,如果内存允许则增加。
  • 多节点部署中的所有节点必须使用相同的 --tensor-parallel-size、--enable-expert-parallel 和模型权重路径配置。

w8a8c8 特有的 --additional-config 字段:

  • "enable_dsa_cp": true:启用 DSA 上下文并行,以加速长上下文 prefill。
  • "enable_balance_scheduling": true:在 v1 调度器中提升输出吞吐并降低 TPOT。当 Prefill-Decode 分离时不建议使用。
  • "fuse_muls_add": true:融合乘加运算。
  • "multistream_overlap_shared_expert": true:在额外的流上重叠 shared-expert 计算。当 "enable_fused_mc2": true 时自动禁用。

5.3 预填充-解码分离

我们希望在多节点环境中展示 GLM-5 的部署指南,采用预填充-解码(PD)分离以获得更好的性能。预填充-解码分离 是指将预填充阶段和解码阶段分布到不同节点上,以提高吞吐量和降低延迟。

In the PD disaggregation scenario, Mooncake is used as the KV cache transfer connector between the prefill and decode nodes. Please refer to KV Cache Pool (Ascend Store) Deployment Guide for the Mooncake configuration.

5.3.1 Prefill-Decode 分离(Ascend 950DT系列产品)

开始之前,请

在每个节点上准备脚本 launch_online_dp.py:

import argparse
import multiprocessing
import os
import subprocess
import sys
def parse_args():
    parser = argparse.ArgumentParser()
    parser.add_argument(
        "--dp-size",
        type=int,
        required=True,
        help="Data parallel size."
    )
    parser.add_argument(
        "--tp-size",
        type=int,
        default=1,
        help="Tensor parallel size."
    )
    parser.add_argument(
        "--dp-size-local",
        type=int,
        default=-1,
        help="Local data parallel size."
    )
    parser.add_argument(
        "--dp-rank-start",
        type=int,
        default=0,
        help="Starting rank for data parallel."
    )
    parser.add_argument(
        "--dp-address",
        type=str,
        required=True,
        help="IP address for data parallel master node."
    )
    parser.add_argument(
        "--dp-rpc-port",
        type=str,
        default=12345,
        help="Port for data parallel master node."
    )
    parser.add_argument(
        "--vllm-start-port",
        type=int,
        default=8000,
        help="Starting port for the engine."
    )
    return parser.parse_args()
args = parse_args()
dp_size = args.dp_size
tp_size = args.tp_size
dp_size_local = args.dp_size_local
if dp_size_local == -1:
    dp_size_local = dp_size
dp_rank_start = args.dp_rank_start
dp_address = args.dp_address
dp_rpc_port = args.dp_rpc_port
vllm_start_port = args.vllm_start_port
def run_command(visible_devices, dp_rank, vllm_engine_port):
    command = [
        "bash",
        "./run_dp_template.sh",
        visible_devices,
        str(vllm_engine_port),
        str(dp_size),
        str(dp_rank),
        dp_address,
        dp_rpc_port,
        str(tp_size),
    ]
    subprocess.run(command, check=True)
if __name__ == "__main__":
    template_path = "./run_dp_template.sh"
    if not os.path.exists(template_path):
        print(f"Template file {template_path} does not exist.")
        sys.exit(1)
    processes = []
    num_cards = dp_size_local * tp_size
    for i in range(dp_size_local):
        dp_rank = dp_rank_start + i
        vllm_engine_port = vllm_start_port + i
        visible_devices = ",".join(str(x) for x in range(i * tp_size, (i + 1) * tp_size))
        process = multiprocessing.Process(target=run_command,
                                        args=(visible_devices, dp_rank,
                                                vllm_engine_port))
        processes.append(process)
        process.start()
    for process in processes:
        process.join()
  1. 在每个节点上准备脚本 run_dp_template.sh。

    1. 预填充节点 0

      #!/usr/bin/env bash
      source /root/.bashrc
      # this obtained through ifconfig
      # nic_name is the network interface name corresponding to local_ip of the current node
      nic_name="xxx"
      local_ip="xxx"
      
      export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=30000
      export HCCL_BUFFSIZE=300
      export HCCL_IF_IP=$local_ip
      export HCCL_SOCKET_IFNAME=$nic_name
      export ASCEND_RT_VISIBLE_DEVICES=$1
      export GLOO_SOCKET_IFNAME=$nic_name
      export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
      export PROMETHEUS_MULTIPROC_DIR=/dev/shm/vllm_metrics && mkdir -p $PROMETHEUS_MULTIPROC_DIR
      export HCCL_DFS_CONFIG="task_exception:off,inconsistent_check:off"
      export HCCL_ALGO=level0:fullmesh
      export ASCEND_LOCAL_COMM_RES='{"version":"1.3"}'
      
      # Ensure the model path matches the directory recorded during download
      vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM5-w4a4 \
          --host 0.0.0.0 \
          --port $2 \
          --data-parallel-size $3 \
          --tensor-parallel-size $7 \
          --max-model-len 135000 \
          --max-num-batched-tokens 8192 \
          --served-model-name glm-5 \
          --gpu-memory-utilization 0.95 \
          --enable-expert-parallel \
          --max-num-seqs 8 \
          --enable-prefix-caching \
          --trust-remote-code \
          --enforce-eager \
          --quantization ascend \
          --enable-auto-tool-choice \
          --tool-call-parser glm47 \
          --reasoning-parser glm45 \
          --kv-cache-dtype fp8 \
          --attention_config.indexer_kv_dtype fp8 \
          --speculative-config '{"num_speculative_tokens": 1, "method": "deepseek_mtp", "enforce_eager": true}' \
          --hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
          --kv-transfer-config \
          '{"kv_connector": "MooncakeConnectorV1",
          "kv_role": "kv_producer",
          "kv_port": "30100",
          "engine_id": "1",
          "kv_connector_extra_config": {
                      "prefill": {
                              "dp_size": 1,
                              "tp_size": 8
                      },
                      "decode": {
                              "dp_size": 16,
                              "tp_size": 1
                      },
                      "ascend_local_comm_res_path": "/etc/hixlep"
              }
          }' \
          --additional-config '{"enable_cpu_binding": "True", "multistream_overlap_shared_expert": "True", "recompute_scheduler_enable": "True", "enable_dsa_cp": true, "enable_flashcomm1": true}'
      
    2. 预填充节点 1

      #!/usr/bin/env bash
      source /root/.bashrc
      # this obtained through ifconfig
      # nic_name is the network interface name corresponding to local_ip of the current node
      nic_name="xxx"
      local_ip="xxx"
      
      export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=30000
      export HCCL_BUFFSIZE=300
      export HCCL_IF_IP=$local_ip
      export HCCL_SOCKET_IFNAME=$nic_name
      export ASCEND_RT_VISIBLE_DEVICES=$1
      export GLOO_SOCKET_IFNAME=$nic_name
      export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
      export PROMETHEUS_MULTIPROC_DIR=/dev/shm/vllm_metrics && mkdir -p $PROMETHEUS_MULTIPROC_DIR
      export HCCL_DFS_CONFIG="task_exception:off,inconsistent_check:off"
      export HCCL_ALGO=level0:fullmesh
      export ASCEND_LOCAL_COMM_RES='{"version":"1.3"}'
      
      # Ensure the model path matches the directory recorded during download
      vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM5-w4a4 \
          --host 0.0.0.0 \
          --port $2 \
          --data-parallel-size $3 \
          --tensor-parallel-size $7 \
          --max-model-len 135000 \
          --max-num-batched-tokens 8192 \
          --served-model-name glm-5 \
          --gpu-memory-utilization 0.95 \
          --enable-expert-parallel \
          --max-num-seqs 8 \
          --enable-prefix-caching \
          --trust-remote-code \
          --enforce-eager \
          --quantization ascend \
          --enable-auto-tool-choice \
          --tool-call-parser glm47 \
          --reasoning-parser glm45 \
          --kv-cache-dtype fp8 \
          --attention_config.indexer_kv_dtype fp8 \
          --speculative-config '{"num_speculative_tokens": 1, "method": "deepseek_mtp", "enforce_eager": true}' \
          --hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
          --kv-transfer-config \
          '{"kv_connector": "MooncakeConnectorV1",
          "kv_role": "kv_producer",
          "kv_port": "30100",
          "engine_id": "1",
          "kv_connector_extra_config": {
                      "prefill": {
                              "dp_size": 1,
                              "tp_size": 8
                      },
                      "decode": {
                              "dp_size": 16,
                              "tp_size": 1
                      },
                      "ascend_local_comm_res_path": "/etc/hixlep"
              }
          }' \
          --additional-config '{"enable_cpu_binding": "True", "multistream_overlap_shared_expert": "True", "recompute_scheduler_enable": "True", "enable_dsa_cp": true, "enable_flashcomm1": true}'
      
    3. 解码节点 0

      #!/usr/bin/env bash
      source /root/.bashrc
      # this obtained through ifconfig
      # nic_name is the network interface name corresponding to local_ip of the current node
      nic_name="xxx"
      local_ip="xxx"
      
      export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=30000
      export HCCL_BUFFSIZE=1200
      export HCCL_IF_IP=$local_ip
      export HCCL_SOCKET_IFNAME=$nic_name
      export ASCEND_RT_VISIBLE_DEVICES=$1
      export GLOO_SOCKET_IFNAME=$nic_name
      export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
      export PROMETHEUS_MULTIPROC_DIR=/dev/shm/vllm_metrics && mkdir -p $PROMETHEUS_MULTIPROC_DIR
      export HCCL_DFS_CONFIG="task_exception:off,inconsistent_check:off"
      export HCCL_ALGO=level0:fullmesh
      export ASCEND_LOCAL_COMM_RES='{"version":"1.3"}'
      
      # Ensure the model path matches the directory recorded during download
      vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM5-w4a4 \
          --host 0.0.0.0 \
          --port $2 \
          --data-parallel-size $3 \
          --data-parallel-rank $4 \
          --data-parallel-address $5 \
          --data-parallel-rpc-port $6 \
          --tensor-parallel-size $7 \
          --max-model-len 135000 \
          --max-num-batched-tokens 240 \
          --served-model-name glm-5 \
          --gpu-memory-utilization 0.95 \
          --enable-expert-parallel \
          --max-num-seqs 60 \
          --enable-prefix-caching \
          --trust-remote-code \
          --quantization ascend \
          --enable-auto-tool-choice \
          --tool-call-parser glm47 \
          --reasoning-parser glm45 \
          --kv-cache-dtype fp8 \
          --attention_config.indexer_kv_dtype fp8 \
          --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
          --speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp", "enforce_eager": true}' \
          --hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
          --kv-transfer-config \
          '{"kv_connector": "MooncakeConnectorV1",
          "kv_role": "kv_consumer",
          "kv_port": "30300",
          "engine_id": "3",
          "kv_connector_extra_config": {
                      "prefill": {
                              "dp_size": 1,
                              "tp_size": 8
                      },
                      "decode": {
                              "dp_size": 16,
                              "tp_size": 1
                      },
                      "ascend_local_comm_res_path": "/etc/hixlep"
              }
          }' \
          --additional-config '{"enable_cpu_binding": "True", "multistream_overlap_shared_expert": "True", "recompute_scheduler_enable": "True", "finegrained_tp_config": {"lmhead_tensor_parallel_size":8}}'
      
    4. 解码节点 1

      #!/usr/bin/env bash
      source /root/.bashrc
      # this obtained through ifconfig
      # nic_name is the network interface name corresponding to local_ip of the current node
      nic_name="xxx"
      local_ip="xxx"
      
      export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=30000
      export HCCL_BUFFSIZE=1200
      export HCCL_IF_IP=$local_ip
      export HCCL_SOCKET_IFNAME=$nic_name
      export ASCEND_RT_VISIBLE_DEVICES=$1
      export GLOO_SOCKET_IFNAME=$nic_name
      export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
      export PROMETHEUS_MULTIPROC_DIR=/dev/shm/vllm_metrics && mkdir -p $PROMETHEUS_MULTIPROC_DIR
      export HCCL_DFS_CONFIG="task_exception:off,inconsistent_check:off"
      export HCCL_ALGO=level0:fullmesh
      export ASCEND_LOCAL_COMM_RES='{"version":"1.3"}'
      
      # Ensure the model path matches the directory recorded during download
      vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM5-w4a4 \
          --host 0.0.0.0 \
          --port $2 \
          --data-parallel-size $3 \
          --data-parallel-rank $4 \
          --data-parallel-address $5 \
          --data-parallel-rpc-port $6 \
          --tensor-parallel-size $7 \
          --max-model-len 202752 \
          --max-num-batched-tokens 240 \
          --served-model-name glm-5 \
          --gpu-memory-utilization 0.95 \
          --enable-expert-parallel \
          --max-num-seqs 60 \
          --enable-prefix-caching \
          --trust-remote-code \
          --quantization ascend \
          --enable-auto-tool-choice \
          --tool-call-parser glm47 \
          --reasoning-parser glm45 \
          --kv-cache-dtype fp8 \
          --attention_config.indexer_kv_dtype fp8 \
          --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
          --speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp", "enforce_eager": true}' \
          --hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
          --kv-transfer-config \
          '{"kv_connector": "MooncakeConnectorV1",
          "kv_role": "kv_consumer",
          "kv_port": "30300",
          "engine_id": "3",
          "kv_connector_extra_config": {
                      "prefill": {
                              "dp_size": 1,
                              "tp_size": 8
                      },
                      "decode": {
                              "dp_size": 16,
                              "tp_size": 1
                      },
                      "ascend_local_comm_res_path": "/etc/hixlep"
              }
          }' \
          --additional-config '{"enable_cpu_binding": "True", "multistream_overlap_shared_expert": "True", "recompute_scheduler_enable": "True", "finegrained_tp_config": {"lmhead_tensor_parallel_size":8}}'
      

准备工作完成后,可以在每个节点上使用以下命令启动服务器:

  1. 预填充节点 0

    # change ip to your own
    python launch_online_dp.py --dp-size 1 --tp-size 8 --dp-size-local 1 --dp-rank-start 0 --dp-address $node_p0_ip --dp-rpc-port 10521 --vllm-start-port 6700
    
  2. 预填充节点 1

    # change ip to your own
    python launch_online_dp.py --dp-size 1 --tp-size 8 --dp-size-local 1 --dp-rank-start 0 --dp-address $node_p1_ip --dp-rpc-port 10521 --vllm-start-port 6700
    
  3. 解码节点 0

    # change ip to your own
    python launch_online_dp.py --dp-size 16 --tp-size 1 --dp-size-local 8 --dp-rank-start 0 --dp-address $node_d0_ip --dp-rpc-port 10523 --vllm-start-port 6721
    
  4. 解码节点 1

    # change ip to your own
    python launch_online_dp.py --dp-size 16 --tp-size 1 --dp-size-local 8 --dp-rank-start 8 --dp-address $node_d0_ip --dp-rpc-port 10523 --vllm-start-port 6721
    

5.3.2 预填充-解码分离(A3 系列)

高吞吐量(198K 上下文)场景已在 4 个 Atlas 800 A3(128GB × 8)上验证:2 个预填充节点(PP2 TP16,78 层划分为 41/37,每个节点一个 PP rank)和 2 个解码节点(DP8 TP4,每个节点 4 个 DP rank)。相同的脚本同时适用于高吞吐量和低延迟场景。

开始之前,请

在每个节点上准备脚本 launch_online_dp.py:

import argparse
import multiprocessing
import os
import subprocess
import sys

def parse_args():
    parser = argparse.ArgumentParser()
    parser.add_argument("--dp-size", type=int, required=True, help="Data parallel size.")
    parser.add_argument("--tp-size", type=int, default=1, help="Tensor parallel size.")
    parser.add_argument("--pp-size", type=int, default=1, help="Pipeline parallel size.")
    parser.add_argument("--dp-size-local", type=int, default=-1, help="Local data parallel size.")
    parser.add_argument("--dp-rank-start", type=int, default=0, help="Starting rank for data parallel.")
    parser.add_argument("--dp-address", type=str, required=True, help="IP address for data parallel master node.")
    parser.add_argument("--dp-rpc-port", type=str, default="12321", help="Port for data parallel master node.")
    parser.add_argument("--vllm-start-port", type=int, default=8000, help="Starting port for the engine.")
    return parser.parse_args()

args = parse_args()
dp_size = args.dp_size
tp_size = args.tp_size
pp_size = args.pp_size
dp_size_local = args.dp_size_local
if dp_size_local == -1:
    dp_size_local = dp_size
dp_rank_start = args.dp_rank_start
dp_address = args.dp_address
dp_rpc_port = args.dp_rpc_port
vllm_start_port = args.vllm_start_port
gpus_per_dp_rank = tp_size * pp_size

def run_command(visible_devices, dp_rank, vllm_engine_port):
    command = [
        "bash",
        "./run_dp_template.sh",
        visible_devices,
        str(vllm_engine_port),
        str(dp_size),
        str(dp_rank),
        dp_address,
        dp_rpc_port,
        str(tp_size),
        str(pp_size),
    ]
    subprocess.run(command, check=True)

if __name__ == "__main__":
    template_path = "./run_dp_template.sh"
    if not os.path.exists(template_path):
        print(f"Template file {template_path} does not exist.")
        sys.exit(1)

    processes = []
    num_cards = dp_size_local * gpus_per_dp_rank

    for i in range(dp_size_local):
        dp_rank = dp_rank_start + i
        vllm_engine_port = vllm_start_port + i
        visible_devices = ",".join(str(x) for x in range(i * gpus_per_dp_rank, (i + 1) * gpus_per_dp_rank))
        process = multiprocessing.Process(
            target=run_command,
            args=(visible_devices, dp_rank, vllm_engine_port)
        )
        processes.append(process)
        process.start()

    for process in processes:
        process.join()
  1. 在每个节点上准备脚本 run_dp_template.sh。

    1. 预填充节点 0

      prefill 脚本通过 node_rank 选择节点:在 prefill 节点 0(PP 主节点,引擎端口 8000)上设置 node_rank=0,在 prefill 节点 1(非主节点,--headless,无 API server)上设置 node_rank=1。

      nic_name="xxxx" # change to your own nic name
      local_ip="xxxx" # change to your own ip
      # pp=2
      # prefill node 0: node_rank=0, prefill node 1: node_rank=1
      node_rank=0
      
      export VLLM_PP_LAYER_PARTITION="41,37"
      export HCCL_BUFFSIZE=400
      export HCCL_IF_IP=$local_ip
      export HCCL_OP_EXPANSION_MODE="AIV"
      export HCCL_SOCKET_IFNAME=$nic_name
      export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/usr/local/lib
      export GLOO_SOCKET_IFNAME=$nic_name
      export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
      
      vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM-5.1-W8A8C8-MTP \
          --host 0.0.0.0 \
          --port 8000 \
          --pipeline-parallel-size 2 \
          --distributed-executor-backend mp \
          --master-addr $local_ip \
          --master-port 7060 \
          --nnodes 2 \
          --node-rank $node_rank \
          --tensor-parallel-size 16 \
          --enable-expert-parallel \
          --speculative-config '{"num_speculative_tokens": 1, "method":"deepseek_mtp","enforce_eager":true}' \
          --seed 1024 \
          --served-model-name glm-5 \
          --max-model-len 202752 \
          --kv-cache-dtype int8 \
          --attention_config.indexer_kv_dtype int8 \
          --additional-config '{"fuse_muls_add": true, "recompute_scheduler_enable": false, "multistream_overlap_shared_expert": true, "enable_dsa_cp": true, "enable_flashcomm1": true, "enable_fused_mc2": true}' \
          --max-num-batched-tokens 16384 \
          --trust-remote-code \
          --enable-prefix-caching \
          --max-num-seqs 64 \
          --quantization ascend \
          --gpu-memory-utilization 0.92 \
          --hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
          --enforce-eager \
          --enable-auto-tool-choice \
          --tool-call-parser glm47 \
          --reasoning-parser glm45 \
          --kv-transfer-config \
          '{"kv_connector": "MooncakeConnectorV1",
          "kv_role": "kv_producer",
          "kv_port": "30000",
          "engine_id": "0",
          "kv_connector_extra_config": {
              "use_ascend_direct": true,
              "prefill": {"dp_size": 1, "pp_size": 2, "tp_size": 16, "pp_layer_partition": "41,37"},
              "decode": {"dp_size": 8, "tp_size": 4}
          }
      }'
      
    2. 预填充节点 1

      nic_name="xxxx" # change to your own nic name
      local_ip="xxxx" # change to your own ip
      # IP of prefill node 0 (the PP master node), must be consistent with the local_ip of prefill node 0
      node_p0_ip="xxxx"
      # pp=2
      # prefill node 0: node_rank=0, prefill node 1: node_rank=1
      node_rank=1
      
      export VLLM_PP_LAYER_PARTITION="41,37"
      export HCCL_BUFFSIZE=400
      export HCCL_IF_IP=$local_ip
      export HCCL_OP_EXPANSION_MODE="AIV"
      export HCCL_SOCKET_IFNAME=$nic_name
      export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/usr/local/lib
      export GLOO_SOCKET_IFNAME=$nic_name
      export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
      
      vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM-5.1-W8A8C8-MTP \
          --host 0.0.0.0 \
          --pipeline-parallel-size 2 \
          --distributed-executor-backend mp \
          --master-addr $node_p0_ip \
          --master-port 7060 \
          --nnodes 2 \
          --node-rank $node_rank \
          --headless \
          --tensor-parallel-size 16 \
          --enable-expert-parallel \
          --speculative-config '{"num_speculative_tokens": 1, "method":"deepseek_mtp","enforce_eager":true}' \
          --seed 1024 \
          --served-model-name glm-5 \
          --max-model-len 202752 \
          --kv-cache-dtype int8 \
          --attention_config.indexer_kv_dtype int8 \
          --additional-config '{"fuse_muls_add": true, "recompute_scheduler_enable": false, "multistream_overlap_shared_expert": true, "enable_dsa_cp": true, "enable_flashcomm1": true, "enable_fused_mc2": true}' \
          --max-num-batched-tokens 16384 \
          --trust-remote-code \
          --enable-prefix-caching \
          --max-num-seqs 64 \
          --quantization ascend \
          --gpu-memory-utilization 0.92 \
          --hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
          --enforce-eager \
          --enable-auto-tool-choice \
          --tool-call-parser glm47 \
          --reasoning-parser glm45 \
          --kv-transfer-config \
          '{"kv_connector": "MooncakeConnectorV1",
          "kv_role": "kv_producer",
          "kv_port": "30000",
          "engine_id": "0",
          "kv_connector_extra_config": {
              "use_ascend_direct": true,
              "prefill": {"dp_size": 1, "pp_size": 2, "tp_size": 16, "pp_layer_partition": "41,37"},
              "decode": {"dp_size": 8, "tp_size": 4}
          }
      }'
      
    3. 解码节点 0(rank 0–3)

      通过位置参数为每个 DP rank 启动一个实例:$1 = 可见设备,$2 = 引擎端口,$3 = 数据并行大小,$4 = 数据并行 rank,$5 = 数据并行地址,$6 = 数据并行 RPC 端口,$7 = 张量并行大小。在解码节点 0 上准备 run_dp_template.sh,内容如下。

      nic_name="xxxx" # change to your own nic name
      local_ip="xxxx" # change to your own ip
      
      # Each DP rank uses an independent engine ID to avoid KV route confusion.
      # $4 = data-parallel-rank. The rank offset within a node is 0 to 3, and the corresponding engine ID is 100 to 103.
      ENGINE_ID=$((100 + $4))
      
      #Mooncake
      
      export HCCL_BUFFSIZE=256
      export HCCL_IF_IP=$local_ip
      export HCCL_OP_EXPANSION_MODE="AIV"
      export HCCL_SOCKET_IFNAME=$nic_name
      export ASCEND_RT_VISIBLE_DEVICES=$1
      export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/usr/local/lib
      export GLOO_SOCKET_IFNAME=$nic_name
      export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
      
      vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM-5.1-W8A8C8-MTP \
          --host 0.0.0.0 \
          --port $2 \
          --data-parallel-size $3 \
          --data-parallel-rank $4 \
          --data-parallel-address $5 \
          --data-parallel-rpc-port $6 \
          --tensor-parallel-size $7 \
          --enable-expert-parallel \
          --speculative-config '{"num_speculative_tokens": 3,  "method":"deepseek_mtp","enforce_eager":true}' \
          --seed 1024 \
          --served-model-name glm-5 \
          --max-model-len 202752 \
          --max-num-batched-tokens 164 \
          --kv-cache-dtype int8 \
          --attention_config.indexer_kv_dtype int8 \
          --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
          --additional-config '{"fuse_muls_add": true, "recompute_scheduler_enable": true, "multistream_overlap_shared_expert": true, "enable_fused_mc2": true, "enable_mlapo": true}' \
          --trust-remote-code \
          --max-num-seqs 32 \
          --gpu-memory-utilization 0.92 \
          --hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
          --async-scheduling \
          --quantization ascend \
          --enable-auto-tool-choice \
          --tool-call-parser glm47 \
          --reasoning-parser glm45 \
          --kv-transfer-config \
          '{"kv_connector": "MooncakeConnectorV1",
          "kv_role": "kv_consumer",
          "kv_port": "30200",
          "engine_id": "'"$ENGINE_ID"'",
          "kv_connector_extra_config": {
              "use_ascend_direct": true,
              "prefill": {"dp_size": 1, "pp_size": 2, "tp_size": 16, "pp_layer_partition": "41,37"},
              "decode": {"dp_size": 8, "tp_size": 4}
          }
      }'
      
    4. 解码节点 1(rank 4–7)

      在解码节点 1 上准备 run_dp_template.sh,内容如下。

      nic_name="xxxx" # change to your own nic name
      local_ip="xxxx" # change to your own ip
      
      # Each DP rank uses an independent engine ID to avoid KV route confusion.
      # $4 = data-parallel-rank. The rank offset within a node is 0 to 3, and the corresponding engine ID is 100 to 103.
      ENGINE_ID=$((100 + $4))
      
      #Mooncake
      
      export HCCL_BUFFSIZE=256
      export HCCL_IF_IP=$local_ip
      export HCCL_OP_EXPANSION_MODE="AIV"
      export HCCL_SOCKET_IFNAME=$nic_name
      export ASCEND_RT_VISIBLE_DEVICES=$1
      export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/usr/local/lib
      export GLOO_SOCKET_IFNAME=$nic_name
      export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
      
      vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/GLM-5.1-W8A8C8-MTP \
          --host 0.0.0.0 \
          --port $2 \
          --data-parallel-size $3 \
          --data-parallel-rank $4 \
          --data-parallel-address $5 \
          --data-parallel-rpc-port $6 \
          --tensor-parallel-size $7 \
          --enable-expert-parallel \
          --speculative-config '{"num_speculative_tokens": 3,  "method":"deepseek_mtp","enforce_eager":true}' \
          --seed 1024 \
          --served-model-name glm-5 \
          --max-model-len 202752 \
          --max-num-batched-tokens 164 \
          --kv-cache-dtype int8 \
          --attention_config.indexer_kv_dtype int8 \
          --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
          --additional-config '{"fuse_muls_add": true, "recompute_scheduler_enable": true, "multistream_overlap_shared_expert": true, "enable_fused_mc2": true, "enable_mlapo": true}' \
          --trust-remote-code \
          --max-num-seqs 32 \
          --gpu-memory-utilization 0.92 \
          --hf-overrides '{"use_index_cache": true, "index_topk_freq": 4}' \
          --async-scheduling \
          --quantization ascend \
          --enable-auto-tool-choice \
          --tool-call-parser glm47 \
          --reasoning-parser glm45 \
          --kv-transfer-config \
          '{"kv_connector": "MooncakeConnectorV1",
          "kv_role": "kv_consumer",
          "kv_port": "30200",
          "engine_id": "'"$ENGINE_ID"'",
          "kv_connector_extra_config": {
              "use_ascend_direct": true,
              "prefill": {"dp_size": 1, "pp_size": 2, "tp_size": 16, "pp_layer_partition": "41,37"},
              "decode": {"dp_size": 8, "tp_size": 4}
          }
      }'
      

准备工作完成后,可以在每个节点上使用以下命令启动服务器:

  1. 预填充节点 0

    bash run_dp_template.sh
    
  2. 预填充节点 1

    bash run_dp_template.sh
    
  3. 解码节点 0

    # change ip to your own
    python launch_online_dp.py --dp-size 8 --tp-size 4 --pp-size 1 --dp-size-local 4 --dp-rank-start 0 --dp-address $node_d0_ip --dp-rpc-port 12321 --vllm-start-port 8000
    
  4. 解码节点 1

    # change ip to your own
    python launch_online_dp.py --dp-size 8 --tp-size 4 --pp-size 1 --dp-size-local 4 --dp-rank-start 4 --dp-address $node_d0_ip --dp-rpc-port 12321 --vllm-start-port 8000
    

注意:

  • 当测试的前缀缓存命中率 > 0 时,在预填充节点上添加 --enable-prefix-caching(如上述脚本所示);当命中率为 0 时,改用 --no-enable-prefix-caching。
  • 在此场景中,预填充节点上的 "recompute_scheduler_enable" 设置为 false,解码节点上设置为 true。

5.3.3 请求转发和关键参数说明

要设置请求转发,请在任意机器上运行以下脚本。您可以在仓库的示例中获取代理程序:load_balance_proxy_server_example.py

Ascend 950DT系列产品:

unset http_proxy
unset https_proxy

python load_balance_proxy_server_example.py \
    --port 8000 \
    --host 0.0.0.0 \
    --prefiller-hosts \
    $node_p0_ip \
    $node_p1_ip \
    --prefiller-ports \
    6700 \
    6700 \
    --decoder-hosts \
    $node_d0_ip \
    $node_d0_ip \
    $node_d0_ip \
    $node_d0_ip \
    $node_d0_ip \
    $node_d0_ip \
    $node_d0_ip \
    $node_d0_ip \
    $node_d1_ip \
    $node_d1_ip \
    $node_d1_ip \
    $node_d1_ip \
    $node_d1_ip \
    $node_d1_ip \
    $node_d1_ip \
    $node_d1_ip \
    --decoder-ports \
    6721 6722 6723 6724 6725 6726 6727 6728 \
    6721 6722 6723 6724 6725 6726 6727 6728

A3 系列:

unset http_proxy
unset https_proxy

python load_balance_proxy_server_example.py \
    --port 8000 \
    --host 0.0.0.0 \
    --prefiller-hosts \
    $node_p0_ip \
    --prefiller-ports \
    8000 \
    --decoder-hosts \
    $node_d0_ip \
    $node_d0_ip \
    $node_d0_ip \
    $node_d0_ip \
    $node_d1_ip \
    $node_d1_ip \
    $node_d1_ip \
    $node_d1_ip \
    --decoder-ports \
    8000 8001 8002 8003 \
    8000 8001 8002 8003

PD 分离部署的关键参数说明:

除了上述单节点和多节点参数外,以下参数专门用于 Prefill-Decode 分离:

Mooncake KV 传输配置(--kv-transfer-config):

  • "kv_connector": "MooncakeConnectorV1":使用 Mooncake 作为 prefill 和 decode 节点之间的 KV 缓存传输连接器。
  • "kv_role": "kv_producer":在 prefill 节点上设置 — 生成 KV 缓存并将其发送到 decode 节点。在 decode 节点上使用 "kv_consumer"。
  • "kv_port":Mooncake KV 传输通信的端口。每个节点组应使用不同的端口范围。
  • "use_ascend_direct": true:启用 Ascend 直接(类似 RDMA)传输 KV 缓存,降低延迟。
  • "prefill" / "decode" 部分:分别指定 prefill 和 decode 节点组的 dp_size 和 tp_size。这些必须与实际部署拓扑匹配。

Prefill 节点特定配置:

  • --additional-config '{"enable_fused_mc2": true}':启用 fused MC2 算子(dispatch_ffn_combine/mega_moe)以优化 MoE 通信。约束条件:dispatch_ffn_combine 仅适用于 w8a8 且 EP≤32;mega_moe 适用于 w8a8/w4a8/bf16 且 EP≤64。两者均与 MTP 和动态 EPLB 不兼容。
  • --additional-config '{"enable_dsa_cp": true}':在预填充节点上启用 DSA 上下文并行以加速长上下文预填充。处理高达 128K token 的提示词时需要此配置。

Decode 节点特定配置:

  • --additional-config '{"enable_mlapo": true}':在 decode 节点上启用 MLA 预处理算子融合,可显著提升 decode 性能。会消耗更多 NPU 内存。在 PD 场景中,仅在 decode 节点上启用 MLAPO。
  • 在 decode 节点上,保持 --max-num-batched-tokens 接近 --max-num-seqs —— decode 每一步每个序列处理一个 token(A3 场景中为 164,Ascend 950DT系列产品场景中为 240,参见上述脚本)。
  • --additional-config '{"recompute_scheduler_enable": true}':启用重计算调度器。当 decode 节点 KV cache 不足时,请求会被送回 prefill 节点进行 KV cache 重计算。在此部署中:decode 节点上为 true;prefill 节点上,Ascend 950DT系列产品场景中为 true,A3 场景中为 false(参见上述脚本)。

常见 PD 环境变量:

  • LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/usr/local/lib:Mooncake库加载所需。

PD 场景中的 MTP:

  • 预填充节点在两种场景下均使用 "num_speculative_tokens": 1(参见上述脚本)。
  • 解码节点在两种场景下均使用 "num_speculative_tokens": 3 以最大化解码吞吐量。
  • 所有预填充和解码节点必须使用相同的 "method": "deepseek_mtp" 和 "enforce_eager": true。

有关上述环境变量的进一步说明和限制,请参阅:envs.py。

6 功能验证

服务器启动后,您可以使用输入提示来查询模型:

curl http://<node0_ip>:<port>/v1/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "glm-5",
        "prompt": "The future of AI is",
        "max_completion_tokens": 15,
        "temperature": 0
    }'

预期结果:

{"id": "chatcmlib-bc44ad093dec79a2", "object": "chat.completion", "created": "1770903266", "model": "glm-5", "choices": [{ "index": 0, "message": {"role": "assistant", "content": "The future of AI is not one thing, but a convergence of several powerful trends.", "annotations": "null", "audio": "null", "function_call": "null", "tool_calls": [], "reasoning": "null"}, "logprobs": "null", "finish_reason": "length", "stop_reason": "null", "token_ids": null}], "service_tier": "null", "system fingerprint": "null", "usage": {"prompt_tokens": 5, "total_tokens": 20, "completion_tokens": 15, "prompt_tokens_details": null}, "prompt_logprobs": "null", "prompt_token_ids": "null", "kv_transfer_params": null}

7 精度评估

7.1 使用 AISBench

  1. 有关详细信息,请参阅 使用 AISBench。

  2. 执行后,您可以获取结果。

8 性能评估

8.1 使用 AISBench

有关详细信息,请参阅 使用 AISBench 进行性能评估。

8.2 使用 vLLM Benchmark

更多详情请参考vllm基准测试。

9 性能调优

9.1 推荐配置

注意:以下配置在特定测试环境中验证,仅供参考。最佳配置取决于最大输入/输出长度、前缀缓存命中率、精度要求和部署机器比例等因素。建议参考调优指南根据实际情况进行调优。

下表提供了Atlas 800 A3上 GLM-5.1-w8a8c8 量化模型的推荐参数配置,涵盖三种部署场景:

测试用例使用 输入/输出 表示法,例如 128k/1k 表示128K输入token和1K输出token;@50%/90% 标记前缀缓存命中率。当前缀缓存命中率 > 0 时,添加 --enable-prefix-caching;当命中率为0时,改为添加 --no-enable-prefix-caching(对于PD场景,这适用于预填充节点)。

9.1.1 表1:详细节点配置

TP/DP 列显示的是部署脚本中配置的每节点数值(一个共置节点承载 4 个 TP4 的 DP rank 使用 16 个 NPU;一个 PD prefill 节点承载 1 个 TP16 的 DP rank 使用 16 个 NPU;一个 PD decode 节点承载 4 个 TP4 的 DP rank 使用 16 个 NPU;一个 Ascend 950DT系列产品节点承载 8 张卡)。198K PD 场景的 prefill 侧使用 PP2 TP16,层划分为 41,37。所有 A3 场景使用 GLM-5.1-w8a8c8 权重;Ascend 950DT系列产品场景使用 GLM-5.1-w4a4 权重。

当使用前缀缓存命中率 > 0 进行测试时,保留 --enable-prefix-caching(如部署脚本中所示);当命中率为 0 时,将其替换为 --no-enable-prefix-caching。

Scenario Weight Version Configuration NPUs TP DP Max Num Seqs Max Num Batched Tokens Max Model Len MTP Spec Num
Dual-Node Co-Located 198K High Throughput (A3) w8a8c8 Dual-Node Co-Located Node (0/1) 16 4 4 6 4096 202752 3
Dual-Node Co-Located 198K Low Latency (A3) w8a8c8 Dual-Node Co-Located Node (0/1) 16 16 1 16 4096 202752 3
PD 198K High Throughput (A3) w8a8c8 PD — Server-P Node (PP2) 16 16 1 64 16384 202752 1
PD 198K High Throughput (A3) w8a8c8 PD — Server-D Node 16 4 4 32 164 202752 3
PD 198K High Throughput (950DT Products) w4a4 PD — Server-P Node (DSA CP 8) 8 8 1 20 8192 202752 1
PD 198K High Throughput (950DT Products) w4a4 PD — Server-D Node 8 1 8 60 240 202752 3
PD 198K Low Latency (950DT Products) w4a4 PD — Server-P Node (DSA CP 8) 8 8 1 20 8192 202752 1
PD 198K Low Latency (950DT Products) w4a4 PD — Server-D Node 8 1 8 60 240 202752 3

9.1.2 表2:需要显式启用的优化

以下优化必须显式启用才能生效。它们适用于A3系列(w8a8c8),具体如下:

Optimization Scenario Enablement Principle (Benefits) Notes
FlashComm_v1 A3预填充节点/共置节点 --additional-config '{"enable_flashcomm1": true}' 将AllReduce拆分为Reduce-Scatter和All-Gather,提升预填充吞吐量并降低通信延迟 当 layer_sharding 包含 o_proj 时不可用
Fused MC2 A3预填充节点 --additional-config '{"enable_fused_mc2": true}' 用 dispatch_ffn_combine/dispatch_gmm_combine_decode 算子替换ALLTOALL+MC2,降低MoE通信开销并提升MoE推理性能 dispatch_ffn_combine 仅适用于w8a8、EP≤32、非MTP、非动态EPLB;与 multistream_overlap_shared_expert 冲突(后者会自动禁用)
MLAPO A3共置高吞吐/PD解码节点 --additional-config '{"enable_mlapo": true}' 融合MLA预处理操作,显著提升解码性能 消耗更多NPU内存;在PD场景中仅在解码节点上启用
DSA CP A3预填充节点;长上下文(≥128K) --additional-config '{"enable_dsa_cp": true}' DSA上下文并行加速长上下文预填充,降低长提示的TTFT 在参考配置中,共置节点和PD预填充节点上启用
Balance Scheduling A3单节点/共置/非PD场景 --additional-config '{"enable_balance_scheduling": true}' 提升v1调度器中的输出吞吐量并降低TPOT TTFT可能劣化;预填充-解码分离时不推荐
Sparse SFA C8 A3(w8a8c8);长上下文预填充 --additional-config '{"enable_sparse_sfa_c8": true}' 稀疏Flash Attention跳过C8量化模型不必要的注意力计算,加速长上下文预填充 v0.23.0中为实验性。在参考配置中,高吞吐和PD场景下启用;低延迟场景下禁用
Sparse LI C8 A3 (w8a8c8) --attention_config.indexer_kv_dtype int8 稀疏注意力优化可减少 C8 量化模型的计算量,提升吞吐量 独立于 enable_sparse_sfa_c8;参考低延迟配置将两者均禁用
重计算调度器 A3解码节点 --additional-config '{"recompute_scheduler_enable": true}' 当解码KV缓存不足时,在预填充节点上重计算KV缓存,避免解码端内存溢出并提高吞吐量 在预填充节点上设置为false
多流重叠共享专家 A3 --additional-config '{"multistream_overlap_shared_expert": true}' 在额外的流上重叠共享专家计算,隐藏其延迟并提高解码性能 当"enable_fused_mc2": true时自动禁用

有关完整的启动命令和详细的参数说明,请参阅在线服务部署中的部署示例和关键参数说明。

9.1.3 表3:性能相关参数调优指南

参数 低延迟 高吞吐 长上下文 描述
--max-num-seqs 较低(16) 较高(6–64) 较高(32–64) 限制并发序列数。较低的值减少调度延迟;较高的值提高吞吐量。
--max-model-len 198K 更长(128K–198K) 最大(198K) 最大上下文长度。必须容纳最长的输入+输出。较大的值消耗更多KV缓存内存。
--max-num-batched-tokens 较低(4096) 预填充较高(4096–16384) 预填充较高(16384);解码节点较小(接近max-num-seqs) 控制每步的批处理大小。较低的值减少每步延迟;较高的值提高预填充吞吐量。
--gpu-memory-utilization 0.92 0.92 0.92 NPU内存比例。本文档中的参考配置使用0.92。如果发生内存溢出,请降低该值。
--enable-chunked-prefill 启用(共置) 启用(共置) 启用(共置) 将长提示拆分为多个块,以防止预填充阻塞解码。PD预填充节点改用--enforce-eager。
num_speculative_tokens(MTP) 3 3 3 MTP推测数量。较高的值提高解码吞吐量,但会消耗草稿模型KV缓存的内存。在参考PD配置中,预填充节点使用1;解码节点使用3。
cudagraph_mode FULL_DECODE_ONLY FULL_DECODE_ONLY(共置/解码节点) FULL_DECODE_ONLY(共置/解码节点) 仅对解码阶段进行图捕获。PD预填充节点改用--enforce-eager。

9.2 调优指南

有关常规性能调优方法,请参阅公开性能调优文档。

有关详细的功能描述和配置选项,请参阅功能指南。

有关环境变量的说明和约束,请参阅envs.py。

10 常见问题解答

  • 常见问题提示:如果遇到问题,请参阅常见问题。

  • 问:如何解决 ValueError: Tokenizer class TokenizersBackend does not exist or is not currently imported?

答:请将 transformers 的版本更新到 5.2.0

  • 问:如何为 GLM-5 启用函数调用?

答:请在 vLLM 启动命令中添加以下配置

--tool-call-parser glm47 \
--reasoning-parser glm45 \
--enable-auto-tool-choice \