跳转至

添加自定义 aclnn 算子

本文档介绍如何向 vllm-ascend 添加自定义 aclnn 算子。

自定义 aclnn 算子在 vllm-ascend 中如何工作?

自定义 aclnn 算子在 vllm-ascend 构建过程中被编译并安装到 vllm_ascend/cann_ops_custom 目录。随后,这些 aclnn 算子被绑定到 torch.ops._C_ascend 模块,使用户能够在 vllm-ascend 的 Python 代码中调用它们。

要启用自定义算子,请使用以下代码:

from vllm_ascend.utils import enable_custom_op

enable_custom_op()

如何添加自定义 aclnn 算子?

  1. csrc 目录下创建一个新的算子文件夹。
  2. 创建 op_hostop_kernel 目录,分别存放主机端和内核端源代码。
  3. csrc/build_aclnn.sh 中为支持的SOC添加构建选项。注意,多个算子应使用 ; 分隔,即 CUSTOM_OPS="op1;op2;op3"
  4. csrc/torch_binding.cpp 中将aclnn算子绑定到 torch.ops._C_ascend 模块。
  5. csrc/torch_binding_meta.cpp 中为算子编写元实现,以便将其捕获到aclgraph中。

成功构建 vllm-ascend 后,即可在 Python 代码中调用自定义 aclnn 算子。