Skip to content

NVFP4 CUDA 13: traceable grouped GEMM

Pre-release
Pre-release

Choose a tag to compare

@S1ro1 S1ro1 released this 09 Sep 19:08
· 1 commit to feat/nvfp4-moe since this release

Built locally on GB200 with CUDA 13 and PyTorch 2.13 from prime-kernels d9a5ceb. Removes torch.compiler.disable from the NVFP4 grouped-GEMM wrapper.

Validation: compiled forward, dgrad, and wgrad match eager exactly for both backward modes; existing GPU numerical tests: 4 passed.

The wheel includes Flash MoE, MXFP8 and NVFP4 kernels. CUDA computation is unchanged. No GitHub Actions build was used.

SHA256: dc73e91ce1ae0dd0174618cb89c4a603b11232dce02a4978d81ac7e979bcb350.