DeepSeek 1M Context Benchmark v1.0.0
This first public release provides a reproducible, sanitized benchmark of deepseek-v4-flash and deepseek-v4-pro across deterministic 32K, 128K, 512K, and approximately 950K prompt tiers. The measured run window was 2026-08-06T20:17:02.706Z to 2026-08-07T00:07:44.737Z.
Highlights
- 344 terminal records: 288 U.S. primary cases, 20 excluded pilot calls, and 36 India network-vantage validation calls.
- Strict exact-match results: Flash 71/144 (49.3%); Pro 81/144 (56.3%).
- U.S. primary end-to-end latency: 12.75 s p50 and 162.29 s p95 (n=288).
- Dated all-role cache-miss cost upper bound: USD 39.913229 with 344/344 usage coverage.
- Sanitized CSV/JSONL data, 11 original charts, generated tables, frozen methodology, and a network-free fixture and grading harness.
Integrity and scope
The release is bound to protocol deepseek-v4-long-context-retrieval-v1.1.0, a pinned source archive, QA report, manifest, and 69 SHA-256 checksums. Raw prompts, provider responses, credentials, cloud identifiers, and private paths are excluded.
These synthetic strict-JSON tasks measure narrow retrieval and synthesis behavior, not general model quality. The India subset is a client network-vantage comparison, not an Indian-user study or server-location claim.
Verify downloaded files with release/checksums.sha256 before analysis.