phase-0
RunPod provisioning, the benchmark harness, and reference-logit capture built and tested (make check green, 26 tests). Real GPU run: 12.75 tokens/sec, 0.355s mean TTFT for deepseek-ai/deepseek-moe-16b-base on a single L40, $0.35 total -- after finding and fixing three real environment bugs in DeepSeek's own remote code and a model-cache location trap. Full account: docs/findings/phase-0/2026-09-14-phase-0-baseline-run.md