NVFP4 CUDA 13 (564904f)
Pre-release
Pre-release
ARM64 CUDA 13 / PyTorch 2.13 kernel wheel built locally on GB200 from 564904f7c1755237e5a97f3f5378d57cfe1736cf.
The NVFP4 operator is prime_rl::grouped_nvfp4_gemm; quantization and dequantization helpers also use the prime_rl namespace with NVFP4-qualified names. The wheel includes the existing Flash MoE and MXFP8 kernels.
Installed-wheel GPU numerical validation: 4 passed. The registered operator schema was checked directly. No GitHub Actions build was used.
SHA256: 7bb6d4db14cb599de97514c78b74aa72de53d77ebd801232f3e1b461429cee10.