NVFP4 CUDA 13: traceable grouped GEMM
Pre-release
Pre-release
·
1 commit
to feat/nvfp4-moe
since this release
Built locally on GB200 with CUDA 13 and PyTorch 2.13 from prime-kernels d9a5ceb. Removes torch.compiler.disable from the NVFP4 grouped-GEMM wrapper.
Validation: compiled forward, dgrad, and wgrad match eager exactly for both backward modes; existing GPU numerical tests: 4 passed.
The wheel includes Flash MoE, MXFP8 and NVFP4 kernels. CUDA computation is unchanged. No GitHub Actions build was used.
SHA256: dc73e91ce1ae0dd0174618cb89c4a603b11232dce02a4978d81ac7e979bcb350.