Skip to content

Feature MatrixΒΆ

This page details the features, kernels, parallelism schemes, and quantization methods currently tested for accuracy and performance.

🚦 Status Legend
  • βœ… Passing: Tested and works as expected. Ready for use.
  • ❌ Failing: Known to be broken or not functional. Help is wanted to fix this!
  • πŸ§ͺ Experimental: Works, but unoptimized or pending community validation.
  • πŸ“ Planned: Not yet implemented, but on the official roadmap.
  • ⛔️ Unplanned: There is no benefit to adding this.
  • ❓ Untested: The functionality exists but has not been recently or thoroughly verified.

Feature Flax Torchax Default
async scheduler βœ… βœ… βœ…
Chunked Prefill βœ… βœ… βœ…
DCN-based P/D disaggregation βœ… βœ… βœ…
KV Cache Offload βœ… βœ… βœ…
LoRA_Torch βœ… βœ… βœ…
Multimodal Inputs βœ… βœ… βœ…
Out-of-tree model support βœ… βœ… βœ…
Prefix Caching βœ… βœ… βœ…
Single Program Multi Data βœ… βœ… βœ…
Speculative Decoding: Eagle3 βœ… βœ… βœ…
Speculative Decoding: Ngram βœ… βœ… βœ…
Speculative Decoding: DFlash βœ… ❓ βœ…
hybrid kv cache ❓ ❓ ❓
multi-host ❓ ❓ ❓
runai_model_streamer_loader ❓ ❓ ❓
sampling_params ❓ ❓ ❓
Step Pooling (Embedding) ❓ ❓ ❓
structured_decoding ❓ ❓ ❓

Feature Flax Torchax Default
async scheduler βœ… βœ… βœ…
Chunked Prefill βœ… βœ… βœ…
DCN-based P/D disaggregation βœ… βœ… βœ…
KV Cache Offload βœ… βœ… βœ…
LoRA_Torch βœ… βœ… βœ…
Multimodal Inputs βœ… βœ… βœ…
Out-of-tree model support βœ… βœ… βœ…
Prefix Caching βœ… βœ… βœ…
Single Program Multi Data βœ… βœ… βœ…
Speculative Decoding: Eagle3 βœ… βœ… βœ…
Speculative Decoding: DFlash βœ… βœ… βœ…
Speculative Decoding: Ngram βœ… βœ… βœ…
hybrid kv cache ❓ ❓ ❓
multi-host ❓ ❓ ❓
runai_model_streamer_loader ❓ ❓ ❓
sampling_params ❓ ❓ ❓
Step Pooling (Embedding) ❓ ❓ ❓
structured_decoding ❓ ❓ ❓

Kernel SupportΒΆ

This table tracks high-level correctness and performance validation for distributed compute kernels.

Feature CorrectnessTest PerformanceTest
Collective Communication Matmul βœ… ❓
MLA ❓ ❓
MoE ❓ ❓
Quantized Attention ❓ ❓
Quantized KV Cache ❓ ❓
Quantized Matmul ❓ ❓
Ragged Paged Attention V3 βœ… βœ…

Microbenchmark Kernel SupportΒΆ

This section outlines the detailed hardware and precision validation for our core microbenchmark kernels.

Category Test W16A16 W8A8 W8A16 W4A4 W4A8 W4A16
Moe Fused MoE ❓ ❓ ❓ ❓ ❓ ❓
gmm ❓ ❓ ❓ ❓ ❓ ❓
Dense All‑gather matmul ❓ ❓ ❓ ❓ ❓ ❓
Attention Generic Ragged Paged
Attention V3
❓ ❓ ❓ ❓ ❓ ❓
MLA ❓ ❓ ❓ ❓ ❓ ❓
Ragged Paged
Attention V3 Head_Dim
64
❓ ❓ ❓ ❓ ❓ ❓

Note: - For attention kernels, W[x]A[y] denotes KV cache as W, A as compute, and x, y as bit precision.

Category Test W16A16 W8A8 W8A16 W4A4 W4A8 W4A16
Moe Fused MoE ❓ ❓ ❓ ❓ ❓ ❓
gmm ❓ ❓ ❓ ❓ ❓ ❓
Dense All‑gather matmul ❓ ❓ ❓ ❓ ❓ ❓
Attention Generic Ragged Paged
Attention V3
❓ ❓ ❓ ❓ ❓ ❓
MLA ❓ ❓ ❓ ❓ ❓ ❓
Ragged Paged
Attention V3 Head_Dim
64
❓ ❓ ❓ ❓ ❓ ❓

Note: - For attention kernels, W[x]A[y] denotes KV cache as W, A as compute, and x, y as bit precision.

Parallelism SupportΒΆ

This table shows the current parallelism support status.

Feature Flax Torchax
Single-host Multi-host Single-host Multi-host
PP βœ… βœ… βœ… βœ…
DP βœ… ❓ βœ… ❓
EP βœ… ❓ βœ… ❓
TP βœ… ❓ βœ… ❓
CP ❓ ❓ ❓ ❓
SP (vote to prioritize) ❓ ❓ ❓ ❓

Feature Flax Torchax
Single-host Multi-host Single-host Multi-host
PP βœ… βœ… βœ… βœ…
DP βœ… ❓ βœ… ❓
EP βœ… ❓ βœ… ❓
TP βœ… ❓ ❌ ❓
CP ❓ ❓ ❓ ❓
SP (vote to prioritize) ❓ ❓ ❓ ❓

Quantization SupportΒΆ

This table shows the current quantization support status.

Checkpoint dtype Method Supported
Hardware Acceleration
Flax Torchax
FP4 W4A16 mxfp4 v7 ❓ ❓
FP8 W8A16 compressed-tensor v7 ❓ ❓
FP8 W8A8 compressed-tensor v7 ❓ ❓
INT4 W4A16 awq v5, v6 ❓ ❓
INT8 W8A8 compressed-tensor v5, v6 ❓ ❓
NVFP4 W4A16 modelopt_fp4 v7 ❓ ❓

Note: - This table only tests checkpoint loading compatibility.

Checkpoint dtype Method Supported
Hardware Acceleration
Flax Torchax
FP4 W4A16 mxfp4 v7 ❓ ❓
FP8 W8A16 compressed-tensor v7 ❓ ❓
FP8 W8A8 compressed-tensor v7 ❓ ❓
INT4 W4A16 awq v5, v6 ❓ ❓
INT8 W8A8 compressed-tensor v5, v6 ❓ ❓
NVFP4 W4A16 modelopt_fp4 v7 ❓ ❓

Note: - This table only tests checkpoint loading compatibility.