NVFP4 CUDA 13 (ec62fa7)
Pre-release
Pre-release
ARM64 CUDA 13 / PyTorch 2.13 kernel wheel built locally on GB200 from ec62fa7c8ce1885a49a21f5ab53b260399105703.
The NVFP4 operator is prime_rl::grouped_nvfp4_gemm; quantization and dequantization helpers also use the prime_rl namespace with NVFP4-qualified names. The wheel includes the existing Flash MoE and MXFP8 kernels.
Installed-wheel GPU numerical validation: 4 passed. The registered operator schema was checked directly. No GitHub Actions build was used.
SHA256: b31d69270f06ba0238e202b2316427cca10aca2576b6fddf84f61cba223c9bcd.