Skip to content

NVFP4 CUDA 13 (564904f)

Pre-release
Pre-release

Choose a tag to compare

@S1ro1 S1ro1 released this 09 Sep 14:59

ARM64 CUDA 13 / PyTorch 2.13 kernel wheel built locally on GB200 from 564904f7c1755237e5a97f3f5378d57cfe1736cf.

The NVFP4 operator is prime_rl::grouped_nvfp4_gemm; quantization and dequantization helpers also use the prime_rl namespace with NVFP4-qualified names. The wheel includes the existing Flash MoE and MXFP8 kernels.

Installed-wheel GPU numerical validation: 4 passed. The registered operator schema was checked directly. No GitHub Actions build was used.

SHA256: 7bb6d4db14cb599de97514c78b74aa72de53d77ebd801232f3e1b461429cee10.