Skip to content

v0.2.0 Fused decode kernels

Latest

Choose a tag to compare

@jaberjaber23 jaberjaber23 released this 29 Jul 11:01

Decode throughput for Kimi-Linear-48B-A3B on one NVIDIA L40S under a hard 32 GiB process cap went from 35.76 to 113.83 tokens per second, a 3.18x gain, with peak reserved memory falling from 29.56 GiB to 27.63 GiB.

The cause was not exotic. Two W4A16 kernels held 77 percent of decode time while running at 4 to 10 percent of the card's memory bandwidth, because both had been written as matrix-matrix products and were being used at decode as matrix-vector products. Rewriting them as real GEMVs took the grouped expert kernel from 8.3 to 51.7 percent of peak.

The fused kernels are not bit-identical to the reference path. Teacher forced they agree on 96.9 percent of next-token choices at 0.0036 nats mean KL, roughly one tenth of the divergence INT4 quantisation itself introduces.

All measurements are one card, one stream, one prompt, greedy decoding. This is not a serving throughput claim and the engine has not been measured on a consumer card.

Full detail in engine/kernels/RESULTS.md and engine/klinear/DECODE-PROFILE.md.