Skip to content

phase-4

@bsreecharanreddy bsreecharanreddy tagged this 21 Sep 15:56
Continuous-batching prefill/decode workers across two real 4-GPU H100
SXM topologies (co-located 4-rank EP, disaggregated 2+2-rank EP with a
real cross-rank KV-cache handoff), both byte-exact against a single-GPU
reference. Three real bugs found and fixed on real hardware, none caught
by CPU-only tests. The measured result was mixed, not clean: TTFT
disaggregation wins at concurrency 4 but loses at concurrency 8, reported
as genuinely inconclusive. Cost: $10.94 of a $40 cap.

Full account: docs/findings/phase-4/2026-09-17-phase-4-disaggregated-prefill-decode-run.md
Assets 2
Loading