What's Changed
- Accept provider total token usage in review gate by @bingran-you in #840
- Add Prime-RL SFT train wrapper by @bingran-you in #842
- fix(integration): unblock expanded review gate by @bingran-you in #841
- Resolve Prime-RL SFT configs from checkout by @bingran-you in #843
- fix(train): support convert from existing Prime-SFT JSONL by @bingran-you in #845
- Add reproducible eval and SFT artifact CLI hooks by @bingran-you in #844
- Harden Prime SFT tool-call validation by @bingran-you in #848
- Redact trajectory secrets before serialization (fix prime-sft convert) by @bingran-you in #849
- Repair Prime SFT conversion from results JSONL by @bingran-you in #850
- Fix OpenRouter and Prime SFT results artifacts by @bingran-you in #851
- Strip OpenHands chat input field before LiteLLM upstream calls by @bingran-you in #852
- Add Prime-RL SFT exposure controls by @bingran-you in #853
- Package local Prime SFT JSONL for Prime-RL by @bingran-you in #854
- Guard Prime-RL SFT reproduction semantics by @bingran-you in #855
- Add Prime-RL custom trainer compatibility controls by @bingran-you in #856
- Add Mobile300 Prime-RL reproduction profile by @bingran-you in #857
- Fix Mobile300 Prime-RL chat-template profile by @bingran-you in #859
- Add Mobile300 Prime-RL tail-window staging by @bingran-you in #860
- Avoid stale Prime-RL local dataset cache by @bingran-you in #861
- Match Mobile300 custom trainer microstep semantics by @bingran-you in #862
- Sync Mobile300 Prime-RL checkpoint step by @bingran-you in #863
- Add Mobile300 Prime-RL token-suffix staging by @bingran-you in #864
- Add Prime-RL sample-mean SFT loss shim by @bingran-you in #865
- Fix Prime-RL custom trainer SFT staging by @bingran-you in #866
- Stub flash-attn imports for Prime-RL SDPA SFT by @bingran-you in #867
- Preserve row-wise Prime-RL SFT exposure by @bingran-you in #869
- Preserve private Prime-SFT tool argument tokens by @bingran-you in #870
- feat(agents): autoload agent plugin packages via benchflow.agents entry points by @Yiminnn in #873
- fix: litellm-proxy health + ACP shims survive parallel/deepseek runs by @Yiminnn in #871
- feat(agents): namespace shorthand + manifest auto-load (#876) by @Yiminnn in #877
- Add native source adapters for MCP benchmarks by @bingran-you in #878
- fix(acp): derive model-via-env from registration data, not agent names by @Yiminnn in #879
- fix(acp): avoid empty agent log placeholders by @bingran-you in #832
- fix(diagnostics): type diagnostic and usage sources by @bingran-you in #858
- fix(litellm): bridge /v1/responses→chat for chat-only OpenAI upstreams by @Yiminnn in #868
- fix(diagnostics): classify core error markers case-insensitively by @bingran-you in #880
- fix(litellm): honor Gemini provider base URL by @bingran-you in #881
- Materialize Toolathlon credential-backed tasks by @bingran-you in #885
- Registry-driven Claude subscription gate + zero-signal heuristic exemption (Anthropic OAuth unlock) by @Yiminnn in #886
- fix(harvey): map provider env to the OpenAI-compatible adapter by @Yiminnn in #888
- Provision Toolathlon service runtime sidecars by @bingran-you in #889
- Fix Toolathlon k8s Daytona runtime by @bingran-you in #890
- Redact sharded worker payload artifacts by @bingran-you in #891
- Stabilize Toolathlon Daytona runtimes by @bingran-you in #892
- Fix Toolathlon Daytona direct runtime by @bingran-you in #893
- Fix Toolathlon Daytona blocker retries by @bingran-you in #894
- Document Python 3.12 CLI install requirement by @bingran-you in #899
- Use noncanonical PTY mode for Daytona ACP by @bingran-you in #896
- Add reusable task runtime primitive by @bingran-you in #902
- Add optional TRL GRPO integration by @bingran-you in #903
- Add HF task snapshot provenance support by @bingran-you in #904
- Add paired eval lift reporter by @bingran-you in #901
- Document BenchFlow GRPO pipeline by @bingran-you in #905
- Skip TRL schema test without transformers by @bingran-you in #906
- Complete TRL rollout artifacts and selective task snapshots by @bingran-you in #907
- Remove BENCHFLOW_SKILL_NUDGE prompt injection by @bingran-you in #908
- Use native data-agent datasets in GRPO runbook by @bingran-you in #909
- fix(acp): restore Gemini on Daytona by @bingran-you in #912
- fix(sandbox): prevent Daytona heartbeat sleep leak by @bingran-you in #910
- Fix OpenHands Azure reasoning effort propagation by @bingran-you in #911
- chore: release v0.6.5 by @bingran-you in #913
Full Changelog: v0.6.4...v0.6.5