Skip to content

phase-1

@bsreecharanreddy bsreecharanreddy tagged this 21 Sep 15:56
A naive and a persistent, cache-aware Triton grouped-GEMM kernel, both
correctness-verified (25/25) on two real GPUs. Swapped into
deepseek-moe-16b-base's 27 MoE layers: ~65-67% faster decode throughput
than DeepSeek's own stock moe_infer, at perfect mutual top-5/top-1 logit
agreement. One honest null result: the persistent kernel's cache-reuse
advantage didn't reproduce at the token count that actually drives
decode throughput. Cost: $0.65.

Full account: docs/findings/phase-1/2026-09-15-phase-1-grouped-gemm-run.md
Assets 2
Loading