Skip to content

Releases: chatdeepai/deepseek-1m-context-benchmark

DeepSeek 1M Context Benchmark v1.0.0

Choose a tag to compare

@chatdeepai chatdeepai released this 07 Aug 11:31

DeepSeek 1M Context Benchmark v1.0.0

This first public release provides a reproducible, sanitized benchmark of deepseek-v4-flash and deepseek-v4-pro across deterministic 32K, 128K, 512K, and approximately 950K prompt tiers. The measured run window was 2026-08-06T20:17:02.706Z to 2026-08-07T00:07:44.737Z.

Highlights

  • 344 terminal records: 288 U.S. primary cases, 20 excluded pilot calls, and 36 India network-vantage validation calls.
  • Strict exact-match results: Flash 71/144 (49.3%); Pro 81/144 (56.3%).
  • U.S. primary end-to-end latency: 12.75 s p50 and 162.29 s p95 (n=288).
  • Dated all-role cache-miss cost upper bound: USD 39.913229 with 344/344 usage coverage.
  • Sanitized CSV/JSONL data, 11 original charts, generated tables, frozen methodology, and a network-free fixture and grading harness.

Integrity and scope

The release is bound to protocol deepseek-v4-long-context-retrieval-v1.1.0, a pinned source archive, QA report, manifest, and 69 SHA-256 checksums. Raw prompts, provider responses, credentials, cloud identifiers, and private paths are excluded.

These synthetic strict-JSON tasks measure narrow retrieval and synthesis behavior, not general model quality. The India subset is a client network-vantage comparison, not an Indian-user study or server-location claim.

Verify downloaded files with release/checksums.sha256 before analysis.