BU Bench V2.1 — 200 tasks
BU Bench V2.1 freezes the current 200-task dataset with the merged task/rubric corrections. Use this tag for reproducible comparisons; the compatible dataset filename remains BU_Bench_V2.enc.
Dataset changes
- Nine tasks have corrections relative to the original release: 014, 028, 029, 048, 053, 088, 113, 163 and 187. These cover search completeness, source-backed aliases, supporting evidence, task wording, changing source counts and blocked property lookups.
- AliExpress 171 and Reuters 185 restore their original rules after the earlier blocked-access edits. Reporting a block does not itself complete an acquisition requirement; Reuters retains its original evidenced-absence outcome.
- All 200 task IDs, item IDs, weights, canaries and 55-task subset membership are unchanged.
- Instructions changed for 014, 028, 029, 088 and 113. Older executions are not executions of those revised instructions.
Judge and runner
The public runner defaults to findings judging and a 60-minute task budget. Judge adapter 2.1.2 compresses/resizes screenshot evidence and preserves revisits; clipping or evidence size does not automatically discard a valid task score. The existing reward-hacking/canary penalty remains. The BrowserCode runner proposal in PR #38 is not part of this release.
Validation and comparability
The existing artifact validator and offline tests pass. Historical results and charts have not been regraded. The newest rules for 014/028/053/163/187 still need saved-evidence judge validation; this release does not claim a new 200-task model result or measured judge accuracy.
The internal content revision remains 2026-09-24-feedback-review-draft-4 to preserve the exact already-validated encrypted bytes. The release version is 2.1; the judge adapter version is separately 2.1.2.
- Full 200-task encrypted dataset
- 55-task subset IDs
- Detailed decisions and remaining work
- Revision metadata and hashes
Dataset SHA-256: 87101f7ebcf3bfbde00091e278741427e906c9b7ccc223892c95e4549704d7c6.
Keep decrypted tasks, rubrics and raw traces private.