Releases: sauravpanda/browser-use-rust
Release list
v0.12.16
Summary
v0.12.16 reverts the v0.12.15 tool-schema slim. The schema slim was intended to lower cached-prefix cost by hiding alias declarations from the default LLM schema, but the WebBench_READ_v5 eval showed it was net-negative: cost/M tokens nearly doubled while accuracy stayed flat.
Why the revert
v0.12.15 vs v0.12.14 per-task cost telemetry on WebBench_READ_v5 (~88 of 198 tasks sampled per run; the 0.12.8-14 lineage tracked tightly in this band):
| Metric | v0.12.14 | v0.12.15 | Δ |
|---|---|---|---|
| Success | 137/198 (69%) | 134/198 (68%) | flat (noise) |
| Cost/M tokens | $0.234 | $0.453 | +94% |
| Cache-hit rate | 62.7% | 19.7% | -43pp |
| Cost/task | $0.0557 | $0.0769 | +38% |
Mechanism: shrinking the cached prefix invalidated Gemini's implicit cross-task cache. Cache-rebuild cost exceeded the per-call schema savings. Same compensation pattern surfaced in the v0.11.x cycle — any change inside the cached prefix forces a cross-task cache rebuild whose cost exceeds the byte savings.
What's in this release
- Pure git revert of merge
70a10cb(PR #22). - Version bumped 0.12.14 → 0.12.16 across
Cargo.toml,Cargo.lock,pyproject.toml. docs/v0.12-plan.mdupdated:- v0.12.15 attempt + result + REJECTED reason.
- v0.12.16 = pure revert (no code lever; substrate sanity-check).
- Next slices pivoted away from cached-prefix trims toward per-step uncached bytes.
Validation
PYTHONPATH=python pytest tests/test_agent_compat.py tests/test_final_answer_guards.py tests/test_blocked_search_guards.py -q— all 99 tests pass.
Expected eval signature
bu-rust 0.12.16 + WebBench_READ_v5 should reproduce v0.12.14's profile (~69% accuracy, ~$0.234/M tokens, ~63% cache hit). A match confirms the v0.12.15 mechanism and clears the substrate for the next experiment.
v0.12.15
Summary
v0.12.15 reduces repeated Gemini prompt overhead by hiding duplicate upstream compatibility alias tool declarations from the default LLM schema.
Aliases such as input_text and press_keys are still accepted by runtime dispatch for compatibility, but the model now sees the canonical tool names by default. Callers that need the old full schema can pass Agent(..., expose_tool_aliases=True).
Cost focus
Local schema measurement dropped the default advertised tool declarations from 62 tools / 33,114 bytes to 44 tools / 22,456 bytes, a 32% schema-byte reduction. This targets cached-prefix cost and first-call fresh prompt size while avoiding canonical tool removal in this release.
Validation
.venv/bin/python -m unittest tests.test_agent_compat tests.test_prompt_metrics.venv/bin/python -m unittest discover -s tests.venv/bin/python -m compileall -q python tests benchcargo check -p bu-pycargo test -p bu-cdp -p bu-dom -p bu-browser.venv/bin/python bench/release_preflight.pygit diff --check
v0.12.14
Summary
- Restores self-validation by default, but only for high-risk final answers.
- Adds
self_validate_risk_onlyso callers can force validate-every-final behavior when needed. - Shortens the validation prompt to reduce the cost of validation turns.
- Updates the v0.12 plan with the 0.12.13 eval result and 0.12.14 adjustment.
Motivation
0.12.13 removed validation turns and reduced cost/steps, but success fell from 140/198 to 132/198 and Incorrect Result failures increased from 16 to 24. Validation was expensive, but it was catching real bad finals. This release restores validation selectively for search/filter/sort/locator tasks, counted top/first/list tasks, recency/current/latest tasks, price/date/availability/fare tasks, and blocked/external-evidence finals.
Dry-run over the 0.12.13 traces estimates validation would apply to about 73 successful step>=3 tasks, versus 102 provisional finalization turns in 0.12.12.
Validation
.venv/bin/python -m unittest discover -s tests.venv/bin/python -m compileall -q python tests benchcargo check -p bu-py.venv/bin/maturin developcargo test -p bu-cdp -p bu-dom -p bu-browser.venv/bin/python bench/release_preflight.py- Installed package reports
browser_use_rs.__version__ == "0.12.14"
v0.12.13
Summary
- Disables global self-validation by default to remove routine
done -> validation -> doneconfirmation turns. - Keeps explicit
self_validate=Truesupport for callers that want the old validation behavior. - Keeps targeted final-answer recovery nudges and top-N/list count checks enabled.
- Updates the v0.12 plan with the 0.12.12 eval result and 0.12.13 adjustment.
Motivation
The 0.12.12 run recovered part of the blocked-site success regression, but trace detail showed 100/198 tasks called done twice and 102/198 tasks had a provisional finalization turn. Those provisional-final tasks averaged about $0.0614 versus $0.0487 for the rest. This release removes that global confirmation round trip while preserving targeted guards.
Validation
.venv/bin/python -m unittest discover -s tests.venv/bin/python -m compileall -q python tests benchcargo check -p bu-py.venv/bin/maturin developcargo test -p bu-cdp -p bu-dom -p bu-browser.venv/bin/python bench/release_preflight.py- Installed package reports
browser_use_rs.__version__ == "0.12.13"
v0.12.12
Summary
- Corrects the 0.12.11 blocked-site fallback regression before the L2 eval.
- Allows bounded external-result completion for static public facts when target-site variants are blocked and the visible result directly answers the task.
- Keeps snippet-only answers rejected for live/current data, prices, availability, bookings, account-gated pages, locators, and site actions.
- Restores softer blocked/search loop force-final thresholds closer to 0.12.10.
- Updates the v0.12 plan with the 0.12.11 eval result and 0.12.12 adjustment.
Validation
.venv/bin/python -m unittest discover -s tests.venv/bin/python -m compileall -q python tests benchcargo check -p bu-py.venv/bin/maturin developcargo test -p bu-cdp -p bu-dom -p bu-browser.venv/bin/python bench/release_preflight.py- Installed package reports
browser_use_rs.__version__ == "0.12.12"
v0.12.11
Summary
- Add a shared blocked-site recovery policy across prompt variants.
- Budget blocked-site recovery to one targeted search fallback, one same-site fallback URL, and one same-site read/extract attempt.
- Remove the validation-path instruction that could force a fresh external search while finalizing.
- Tighten existing bot-block/search-fallback loop windows so blocked/search-engine tails stop earlier.
- Mark search-engine challenge responses as consuming the search fallback budget.
Validation
.venv/bin/maturin develop.venv/bin/python -m unittest tests.test_blocked_search_guards tests.test_final_answer_guards tests.test_batch_guard_handling.venv/bin/python -m unittest discover -s tests.venv/bin/python -m compileall -q python tests benchcargo check -p bu-pycargo test -p bu-cdp -p bu-dom -p bu-browser.venv/bin/python bench/release_preflight.py- version import check for
0.12.11
v0.12.10
Changes
- Reuse the previous full PAGE_STATE prompt when URL + cleaned LLM-facing DOM are unchanged.
- Append a small PAGE_STATE_UNCHANGED marker instead of resending another screenshot/DOM blob, allowing provider prompt caching to reuse the retained full state.
- Force a full refresh after 3 reused steps by default.
- Add
state_cache_max_reuse_steps,BROWSER_USE_RS_STATE_CACHE_MAX_REUSE_STEPS, andBU_RS_STATE_CACHE_MAX_REUSE_STEPS; set to0to disable. - Add prompt metadata for page-state bytes, page-state image bytes, reuse marker count, and reuse count.
Notes
- The cache key uses cleaned browser state: URL + rendered LLM DOM. It does not use raw DOM, and it ignores exact screenshot bytes to avoid animation/ad churn defeating cache reuse.
- Full PAGE_STATE is not reused when the retained full state contained transient
<read_state>.
Validation
.venv/bin/maturin develop.venv/bin/python -m unittest discover -s tests.venv/bin/python -m compileall -q python tests benchcargo check -p bu-pycargo test -p bu-cdp -p bu-dom -p bu-browser.venv/bin/python bench/release_preflight.py
v0.12.9
Changes
- Cap the auto-injected DOM page state at 64 KiB by default to reduce repeated text prompt cost on large pages.
- Add
dom_max_bytes,BROWSER_USE_RS_DOM_MAX_BYTES, andBU_RS_DOM_MAX_BYTEScontrols; set to0to disable the cap. - Filter valid action indices to the DOM rows actually shown to the model, so omitted
[N]values cannot be used accidentally. - Add DOM cap metadata for follow-up evaluation analysis.
Validation
.venv/bin/maturin develop.venv/bin/python -m unittest discover -s tests.venv/bin/python -m compileall -q python tests benchcargo check -p bu-pycargo test -p bu-cdp -p bu-dom -p bu-browser.venv/bin/python bench/release_preflight.py
v0.12.8
Summary
- Reduce Gemini screenshot/media token cost by defaulting ChatGoogle media inputs to low resolution.
- Honor
Agent(..., vision_detail_level=...)by forwarding it to providers with media-resolution support. - Honor
images_per_step=0by omitting screenshot image parts from LLM prompts while keeping DOM state. - Add modality token details to usage logs and
ChatInvokeUsage.model_dump()for image/text cost attribution.
Verification
.venv/bin/python -m unittest discover -s tests -q.venv/bin/python -m compileall -q python/browser_use_rs tests benchcargo check -p bu-pycargo test -p bu-cdp -p bu-dom -p bu-browserBROWSER_USE_RS_DISABLE_DOTENV=1 .venv/bin/python bench/release_preflight.pygit diff --check
v0.12.7
- Restore conservative default web_search behavior after the 0.12.6 eval regression: search opens the SERP and returns immediately by default.\n- Keep the richer SERP snippet extraction available behind BROWSER_USE_RS_WEB_SEARCH_SNIPPETS=1 for isolated testing.\n- Retain the Anthropic Haiku fix so effort is only sent when thinking is active.