Skip to content

phase-5a

@bsreecharanreddy bsreecharanreddy tagged this 21 Sep 15:56
Self-computed, per-output-channel int8 quantization extending the Triton
kernel itself (not sourced from bitsandbytes/AWQ). Kernel-level
correctness gate passed on a real L40 (15/15 int8, 25/25 bf16). A real
OOM bug (quantizing without freeing the original bf16 weights) found and
fixed mid-session; whole-branch review caught two more real precision
defects before merge. Measured: +70.9% over stock (a throughput tie with
the bf16 kernel it's built on, as expected for weight-only
quantization), 49.89% expert-weight memory reduction, perfect
model-level agreement. Cost: $1.95 of a $5 cap.

Full account: docs/findings/phase-5a/2026-09-17-phase-5a-quantization-run.md
Assets 2
Loading