vllm.model_executor.kernels.mhc.aiter ¶
Functions:
-
mhc_fused_post_pre_aiter–Fused mHC post + next mHC pre on ROCm via AITER.
-
mhc_pre_aiter–Forward pass for mHC pre block.
-
mhc_pre_delayed_aiter–MHC pre with the pre-mix carried in from the previous sublayer.
mhc_fused_post_pre_aiter(x, residual, post_layer_mix, comb_res_mix, fn, hc_scale, hc_base, rms_eps, hc_pre_eps, hc_sinkhorn_eps, hc_post_mult_value, sinkhorn_repeat, n_splits=1, tile_n=1, norm_weight=None, norm_eps=0.0) ¶
Fused mHC post + next mHC pre on ROCm via AITER.
Returns residual_cur, post_mix_cur, comb_mix_cur, layer_input_cur.
Source code in vllm/model_executor/kernels/mhc/aiter.py
mhc_pre_aiter(residual, fn, hc_scale, hc_base, rms_eps, hc_pre_eps, hc_sinkhorn_eps, hc_post_mult_value, sinkhorn_repeat, n_splits=1, norm_weight=None, norm_eps=0.0) ¶
Forward pass for mHC pre block.
Parameters:
-
(residual¶Tensor) –shape (..., hc_mult, hidden_size), dtype torch.bfloat16
-
(fn¶Tensor) –shape (hc_mult3, hc_mult * hidden_size), dtype torch.float32
-
(hc_scale¶Tensor) –shape (3,), dtype torch.float32
-
(hc_base¶Tensor) –shape (hc_mult3,), dtype torch.float32
-
(rms_eps¶float) –RMS normalization epsilon
-
(hc_pre_eps¶float) –pre-mix epsilon
-
(hc_sinkhorn_eps¶float) –sinkhorn epsilon
-
(hc_post_mult_value¶float) –post-mix multiplier value
-
(sinkhorn_repeat¶int) –number of sinkhorn iterations
-
(n_splits¶int, default:1) –split-k factor;
-
(norm_weight¶Tensor | None, default:None) –optional RMSNorm weight fused into the pre kernel
-
(norm_eps¶float, default:0.0) –epsilon for the fused RMSNorm when norm_weight is set
Returns:
-
post_mix(Tensor) –shape (..., hc_mult), dtype torch.float32
-
comb_mix(Tensor) –shape (..., hc_mult, hc_mult), dtype torch.float32
-
layer_input(Tensor) –shape (..., hidden_size), dtype torch.bfloat16
Source code in vllm/model_executor/kernels/mhc/aiter.py
mhc_pre_delayed_aiter(residual, fn, hc_scale, hc_base, rms_eps, hc_pre_eps, hc_sinkhorn_eps, hc_post_mult_value, sinkhorn_repeat, pre_mix=None, sublayer_out=None, post_layer_mix=None, comb_res_mix=None, residual_out=None) ¶
MHC pre with the pre-mix carried in from the previous sublayer.
Matches mhc_pre_delayed_torch: the stream collapse uses pre_mix rather than the gate computed here, and that gate is returned as the pre-mix for the next sublayer seam.
Parameters:
-
(residual¶Tensor) –shape (..., hc_mult, hidden_size), dtype torch.bfloat16
-
(fn¶Tensor) –shape (hc_mult3, hc_mult * hidden_size), dtype torch.float32
-
(hc_scale¶Tensor) –shape (3,), dtype torch.float32
-
(hc_base¶Tensor) –shape (hc_mult3,), dtype torch.float32
-
(rms_eps¶float) –RMS normalization epsilon
-
(hc_pre_eps¶float) –pre-mix epsilon
-
(hc_sinkhorn_eps¶float) –sinkhorn epsilon
-
(hc_post_mult_value¶float) –post-mix multiplier value
-
(sinkhorn_repeat¶int) –number of sinkhorn iterations
-
(pre_mix¶Tensor | None, default:None) –shape (..., hc_mult) from the previous sublayer, or None at model entry to select residual stream zero.
-
(sublayer_out¶Tensor | None, default:None) –attention or FFN output, shape (..., hidden_size). When given with the two mixes below, the preceding post block is applied here so AITER can fold it into the pre projection.
-
(post_layer_mix¶Tensor | None, default:None) –shape (..., hc_mult, 1), post gate for that block.
-
(comb_res_mix¶Tensor | None, default:None) –shape (..., hc_mult, hc_mult), residual comb for it.
-
(residual_out¶Tensor | None, default:None) –shape (..., hc_mult, hidden_size), written with the post block's new residual. Required exactly when the post is folded in.
Returns:
-
post_mix(Tensor) –shape (..., hc_mult, 1), dtype torch.float32
-
comb_mix(Tensor) –shape (..., hc_mult, hc_mult), dtype torch.float32
-
layer_input(Tensor) –shape (..., hidden_size), dtype torch.bfloat16
-
next_pre_mix(Tensor) –shape (..., hc_mult), dtype torch.float32