Backward pass CUDA kernels for fused GEMM+Bias+GeLU on SM75 (Turing). Float4 vectorized WMMA kernels validated against PyTorch autograd. 3.1x faster than autograd at M=1024. 24/24 tests passing.
-
Updated
Jul 27, 2026 - Python
Backward pass CUDA kernels for fused GEMM+Bias+GeLU on SM75 (Turing). Float4 vectorized WMMA kernels validated against PyTorch autograd. 3.1x faster than autograd at M=1024. 24/24 tests passing.
Compiler MVP that detects Transformer fusion patterns, generates optimized CUDA kernels with WMMA Tensor Cores, and executes them on real GPU hardware — 10.5 TFLOPs on RTX 2070, correctness validated against PyTorch.
FlashAttention v1 forward pass in CUDA for NVIDIA Turing (SM75)
Add a description, image, and links to the sm75 topic page so that developers can more easily learn about it.
To associate your repository with the sm75 topic, visit your repo's landing page and select "manage topics."