kernels a continuously expanding collection of high-performance ML-centric (so far) CUDA kernels some of these kernels have their naive + optimized versions, with everything in-between, and some require CUTLASS CuTe ops. supported: conv2d binary elementwise fp32fp32 GEMM 2d transpose