Skip to content

benchflow 0.7.6

Latest

Choose a tag to compare

@github-actions github-actions released this 04 Sep 20:38
· 1 commit to main since this release
28f88bb

What's Changed

  • Bump BenchFlow to 0.7.6.dev0 by @bingran-you in #1042
  • fix(acp): modernize codex-acp dispatch — shim 1.6.0, effort via model id by @bingran-you in #1044
  • Preserve healthy native ACP subscription results by @kywch in #1049
  • Add --publish-bucket and --eval-results-* flags to bench eval run by @NathanHB in #1035
  • feat(viewer): reviewer-grade trajectory pages, multi-run browsing, and hf:// dataset sources by @ljr145733 in #1034
  • fix(deps): update RestrictedPython to 8.3 by @kywch in #1070
  • Cap OpenClaw Claude output tokens by @kywch in #1075
  • fix(cli): reject --trials > 1 without --matrix by @Benjamin-eecs in #1064
  • Bump claude-agent-acp pin 0.40.0 -> 0.73.0 so claude-fable-5-1 can run by @yuchenlwu in #1086
  • fix(sandbox): use the verifier-setup budget for hardening execs by @Benjamin-eecs in #1062
  • fix(eval): re-run infra-retryable verifier-errored tasks on resume by @Benjamin-eecs in #1063
  • fix(acp): cap how long a pending tool call defers the idle watchdog by @Benjamin-eecs in #1066
  • Fix Gemini ACP model IDs with Google provider prefixes by @kywch in #1030
  • fix(judge): parse native edit targets by @kywch in #1069
  • Honor Codex proxy-owned model selection by @kywch in #1076
  • fix(acp): preserve usage and terminal evidence on timeout by @kywch in #1080
  • Add Z.ai Coding Plan routing by @kywch in #1074
  • fix(litellm): provision Vertex sandbox runtime by @kywch in #985
  • fix: guard _classify_completed_outcomes against non-dict rewards by @VaisakhiMishra in #1054
  • Release BenchFlow 0.7.6 by @bingran-you in #1103

New Contributors

Full Changelog: v0.7.5...v0.7.6