A naive and a persistent, cache-aware Triton grouped-GEMM kernel, both
correctness-verified (25/25) on two real GPUs. Swapped into
deepseek-moe-16b-base's 27 MoE layers: ~65-67% faster decode throughput
than DeepSeek's own stock moe_infer, at perfect mutual top-5/top-1 logit
agreement. One honest null result: the persistent kernel's cache-reuse
advantage didn't reproduce at the token count that actually drives
decode throughput. Cost: $0.65.
Full account: docs/findings/phase-1/2026-09-15-phase-1-grouped-gemm-run.md