This repository was archived by the owner on Feb 3, 2026. It is now read-only.
v0.22.0
Changes
- Simplify agent modes from three (text/visual/full) to two (dom/pixel)
- dom mode (default): Uses element IDs for interaction, includes DOM + screenshots
- pixel mode: Uses screen coordinates, includes screenshots only
- Move ScrollDocumentTool to pixel mode only
- Both modes now get screenshots for better LLM context