DreamBench-SWE v2.1.0 is the additive successor release for Paper A. It publishes the current manuscript and the separately preregistered external-systems audit while preserving the original v2.0.5 tag and every v2.0.5 asset unchanged.
The successor result is deliberately conservative: external discriminant validity was not statistically demonstrated. This release does not claim superiority, equivalence, or broad product-level generalization. Proposed external conditions that failed the frozen pre-evaluation conformance gate were not benchmark-evaluated.
Release assets:
- Sanitized self-contained evidence artifact:
e8b8d27c7d532aee477419fd90904d93a734378beec4d46d7a06a5fe29463c2b - Paper PDF:
f0201fdc6a380fcf045fd0054340e4b6fbdc86afede1d96a54df7ace9db0622f - Flat arXiv source:
f07216c7c3cf4ad582d1f504c757173c20adcec96561b51496b97860eee525c8 - Manifest:
34c8fd8ecc5d554029811e5ca7b4e5cbf3fc7460c0d969cb932b92b8e2b5ed07 - Detached release ledger:
9e8654e8875c4466febf02addfb6afd7e0adecf9117434aa144e5c193f951090
The manifest binds the frozen evaluation source, launcher source, and release source. The evidence artifact ships a public verifier and selected regression tests; raw hosted-model logs, credentials, private analyzer inputs, hidden oracles, and local operational paths are excluded.