Skip to content

Releases: henryqin1997/statem

StateM DeepSeek V4 Flash policy-v9 — Terminal-Bench 2.1 artifacts

Choose a tag to compare

This release publishes the source and redacted evaluation evidence for the
StateM DeepSeek V4 Flash policy-v9 Terminal-Bench 2.1 run discussed in the
StateM paper.

Assets

  • statem-deepseek-v4-flash-policy9-tb21-source-exact-20260818.tar.gz
    contains the exact 54-file task-injected source set. Its verifier confirms
    manifest SHA-256
    4b8aa059f87eee1a7b6f7cc3af8ef94ec628c68a48f56adca62069f44189d075.
  • statem-deepseek-v4-flash-policy9-tb21-reproduction-kit-20260818.tar.gz
    adds a separately identified host-side DeepSeek bridge, frozen control
    plane, credential-free provider template, and tested Harbor dry-run guide.
  • statem-deepseek-v4-flash-0731-policy9-88task-k5-public-redacted-20260813.tar.gz
    contains 440 redacted ATIF trajectories plus StateM states, routes, checks,
    receipts, and result metadata.
  • SHA256SUMS authenticates the three archives.

Result boundary

The public artifact covers 88 tasks with five trials per task and excludes
gpt2-codegolf. It records 392/440 raw passes (89.09%). The paper's standard
89-task denominator reports the same 392 passes as 392/445 (88.09%).

Provider credentials, provider configuration paths, raw session trees,
duplicate Codex output, verifier artifacts, and job/trial logs are not
published. Credential-shaped strings from the public sanitize-git-repo
fixture and sandbox-generated certificate keys are retained as benchmark
evidence and documented inside the artifact.

Citation

@misc{qin2026statemreaching953raw,
  title         = {StateM: Reaching 95.3\% Raw Accuracy, or a \$15 Frontier Run,
                   on Terminal-Bench 2.1 via Harness Scaling},
  author        = {Ziheng Qin and Yaxin Lu and Zhangyang Atlas Wang and Kai Wang},
  year          = {2026},
  eprint        = {2608.15089},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2608.15089}
}