Skip to content

NVFP4 4over6 CUDA 13 ARM64 (436719d)

Pre-release
Pre-release

Choose a tag to compare

@S1ro1 S1ro1 released this 10 Sep 00:52

ARM64 / Python 3.12 / CUDA 13 / torch 2.13.0 wheel built locally on GB200 from 436719d. Contains flash_moe, mxfp8_moe, and nvfp4_moe.

Adds optional four_over_six=True to NVFP4 weight/activation quantizers and grouped GEMM. Defaults to False. Enabled arithmetic matches FlashInfer's default 4/6 recipe: 448 normalization, MAE, strict error scoring and default approximate candidate arithmetic.

SHA256: e8311b13da602def10c4cc88ac6abad2c8d5090799d04a14a3c89421257a0d68

The installed wheel passed 8 numerical tests, 15 exact comparisons against FlashInfer/vLLM defaults, compiled forward/backward equality, and a Prime-RL MoE forward/backward smoke with non-zero router/expert gradients. All registered kernels report available. No GitHub Actions build was used.