Releases: henryqin1997/statem
Releases · henryqin1997/statem
Release list
StateM DeepSeek V4 Flash policy-v9 — Terminal-Bench 2.1 artifacts
This release publishes the source and redacted evaluation evidence for the
StateM DeepSeek V4 Flash policy-v9 Terminal-Bench 2.1 run discussed in the
StateM paper.
Assets
statem-deepseek-v4-flash-policy9-tb21-source-exact-20260818.tar.gz
contains the exact 54-file task-injected source set. Its verifier confirms
manifest SHA-256
4b8aa059f87eee1a7b6f7cc3af8ef94ec628c68a48f56adca62069f44189d075.statem-deepseek-v4-flash-policy9-tb21-reproduction-kit-20260818.tar.gz
adds a separately identified host-side DeepSeek bridge, frozen control
plane, credential-free provider template, and tested Harbor dry-run guide.statem-deepseek-v4-flash-0731-policy9-88task-k5-public-redacted-20260813.tar.gz
contains 440 redacted ATIF trajectories plus StateM states, routes, checks,
receipts, and result metadata.SHA256SUMSauthenticates the three archives.
Result boundary
The public artifact covers 88 tasks with five trials per task and excludes
gpt2-codegolf. It records 392/440 raw passes (89.09%). The paper's standard
89-task denominator reports the same 392 passes as 392/445 (88.09%).
Provider credentials, provider configuration paths, raw session trees,
duplicate Codex output, verifier artifacts, and job/trial logs are not
published. Credential-shaped strings from the public sanitize-git-repo
fixture and sandbox-generated certificate keys are retained as benchmark
evidence and documented inside the artifact.
Citation
@misc{qin2026statemreaching953raw,
title = {StateM: Reaching 95.3\% Raw Accuracy, or a \$15 Frontier Run,
on Terminal-Bench 2.1 via Harness Scaling},
author = {Ziheng Qin and Yaxin Lu and Zhangyang Atlas Wang and Kai Wang},
year = {2026},
eprint = {2608.15089},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2608.15089}
}