Skip to content

phase-6

@bsreecharanreddy bsreecharanreddy tagged this 21 Sep 15:56
The at-scale correctness gate Phase 5b owed: naive, persistent, and int8
kernels all agree with the stock reference at 1,036 positions (97.1-97.3%
top-1 agreement), zero large-gap disagreements. The kernel race is a
real, disclosed loss for dispatch -- vLLM 0.29.0 beats dispatch's kernels
at every one of 28 measured (precision, token count, distribution)
combinations on identical hardware, largest gap at 1 token (2.4-4.1x
slower), narrowing but never closing by 128+ tokens. dispatch does beat
SGLang specifically at 128 and 512 tokens, both precisions. vLLM and
SGLang each served the real model correctly across concurrency
1/4/16/64; dispatch has no served, batched engine of its own to place in
that table. Six real fixes shipped mid-session, all recorded in the
runbook. Cost: $4.5448 of a $10 cap.

Full account: docs/findings/phase-6/2026-09-20-phase-6-final-benchmark-run.md
Runbook: docs/runbooks/phase-6-final-benchmark.md
Assets 2
Loading