What's Changed
- Bump BenchFlow to 0.7.6.dev0 by @bingran-you in #1042
- fix(acp): modernize codex-acp dispatch — shim 1.6.0, effort via model id by @bingran-you in #1044
- Preserve healthy native ACP subscription results by @kywch in #1049
- Add --publish-bucket and --eval-results-* flags to bench eval run by @NathanHB in #1035
- feat(viewer): reviewer-grade trajectory pages, multi-run browsing, and hf:// dataset sources by @ljr145733 in #1034
- fix(deps): update RestrictedPython to 8.3 by @kywch in #1070
- Cap OpenClaw Claude output tokens by @kywch in #1075
- fix(cli): reject --trials > 1 without --matrix by @Benjamin-eecs in #1064
- Bump claude-agent-acp pin 0.40.0 -> 0.73.0 so claude-fable-5-1 can run by @yuchenlwu in #1086
- fix(sandbox): use the verifier-setup budget for hardening execs by @Benjamin-eecs in #1062
- fix(eval): re-run infra-retryable verifier-errored tasks on resume by @Benjamin-eecs in #1063
- fix(acp): cap how long a pending tool call defers the idle watchdog by @Benjamin-eecs in #1066
- Fix Gemini ACP model IDs with Google provider prefixes by @kywch in #1030
- fix(judge): parse native edit targets by @kywch in #1069
- Honor Codex proxy-owned model selection by @kywch in #1076
- fix(acp): preserve usage and terminal evidence on timeout by @kywch in #1080
- Add Z.ai Coding Plan routing by @kywch in #1074
- fix(litellm): provision Vertex sandbox runtime by @kywch in #985
- fix: guard _classify_completed_outcomes against non-dict rewards by @VaisakhiMishra in #1054
- Release BenchFlow 0.7.6 by @bingran-you in #1103
New Contributors
- @NathanHB made their first contribution in #1035
- @ljr145733 made their first contribution in #1034
- @Benjamin-eecs made their first contribution in #1064
- @yuchenlwu made their first contribution in #1086
- @VaisakhiMishra made their first contribution in #1054
Full Changelog: v0.7.5...v0.7.6