A shared propose/verify/accept/rollback loop with two drafters (a 7B
draft model, and model-free prompt-lookup) on top of Phase 5a's int8
target. The first GPU session's numbers were withdrawn by the final
review (degenerate baseline); root-caused to a transformers==5.17.0 bug
leaving DeepSeek's remote-code RoPE inv_freq buffer uninitialized,
poisoning attention with NaN on GPU. Real re-run on an A40: baseline
25.18 tok/s; only prompt-lookup beats it (34.4-50.9 tok/s across k=1-8;
draft-model stays below baseline at every k). Not every config is
byte-exact against the baseline -- traced to a genuine near-tied logit
under int8 precision, not a logic bug. Cost: $12.26 across both
sessions, against a $10 cap raised to $20 with disclosure.
Full account: docs/findings/phase-5b/2026-09-18-phase-5b-speculative-decoding-run.md