v1.9.0
Added support for experimental critic model
critic-demo.mp4
What's Changed
- Propagate Datadog conversation logging flag to evaluation dispatch by @simonrosenberg in #1703
- Release v1.8.2 by @all-hands-bot in #1722
- Add swebenchmultimodal support to run-eval workflow by @juanmichelini in #1659
- feat: add multiswebench support to run-eval workflow by @juanmichelini in #1693
- feat: load skills from agent server by @hieptl in #1729
- Add GPT-5.2 high reasoning to model configurations by @neubig in #1735
- Replace mocked MCP tests with real integration tests by @neubig in #1678
- feat(plugin): Add Plugin.fetch() for remote plugin fetching and caching by @jpshackelford in #1647
- Fix event ordering in RemoteEventsList by inserting events sorted by timestamp by @neubig in #1737
- fix: move async operations outside sync lock in event subscription by @neubig in #1732
- Remove unsupported models by @simonrosenberg in #1734
- refactor(hooks): add typed fields to HookConfig for better type safety by @xingyaoww in #1726
- Add Marketplace datamodel for marketplace.json support by @neubig in #1744
- Remove push_to_index option from run-eval.yml workflow by @juanmichelini in #1750
- feat: add Gemini 3 models to reasoning effort supported list by @Chesars in #1752
- Fix DiscriminatedUnionEnvParser to use single parser when only one kind exists by @tofarr in #1741
- Add GitHub repository URLs to PyPI package metadata by @xingyaoww in #1753
- Fix cache tag truncation with ports and latest suffix by @simonrosenberg in #1626
- docs: rewrite CONTRIBUTING around SDK architecture by @enyst in #1755
- Fix CORS to allow DOCKER_HOST_ADDR for remote browser access by @shanemort1982 in #1466
- refactor(examples): use .from_dict() for hooks example by @xingyaoww in #1746
- fix: Add JSON formatter for uvicorn access logs by @neubig in #1733
- Fix unhandled ConversationRunError causing non-JSON logs by @neubig in #1680
- remove gpt-5-mini from list of models to evaluate for openhands index by @simonrosenberg in #1766
- Add API-Based Critic for Real-Time Agent Action Evaluation (Experimental) by @xingyaoww in #1269
- feat: Add .pr/ directory convention and auto-cleanup workflow by @jpshackelford in #1764
- feat(delegate): Add create_sub_visualizer method for custom sub-agent visualization by @malhotra5 in #1767
- Add minimax/MiniMax-M2.1 to resolve_model_config.py by @juanmichelini in #1757
- feat: Support full class names in DiscriminatedUnionEnvParser by @tofarr in #1768
- Update idle time on bash, git, and file operations by @neubig in #1770
- Add set_security_analyzer abstract method to BaseConversation by @malhotra5 in #1772
New Contributors
- @shanemort1982 made their first contribution in #1466
Full Changelog: v1.8.2...v1.9.0