vllm.model_executor.layers.quantization.utils.humming.priority
¶
Humming kernel selection priority.
Functions:
-
prefers_humming–Whether Humming outranks Marlin. A missing compute capability uses the
-
prioritize_humming–Move Humming directly ahead of Marlin where Humming is preferred.
prefers_humming(compute_capability=None)
¶
Whether Humming outranks Marlin. A missing compute capability uses the current device.
Source code in vllm/model_executor/layers/quantization/utils/humming/priority.py
prioritize_humming(kernels, compute_capability=None)
¶
Move Humming directly ahead of Marlin where Humming is preferred.
Every other entry keeps its relative order and the input list is not modified. Match kernel class names or MoE backend enum names.