Skip to content

Releases: sauravpanda/browser-use-rust

v0.12.16

Choose a tag to compare

@sauravpanda sauravpanda released this 28 May 19:33
608f75b

Summary

v0.12.16 reverts the v0.12.15 tool-schema slim. The schema slim was intended to lower cached-prefix cost by hiding alias declarations from the default LLM schema, but the WebBench_READ_v5 eval showed it was net-negative: cost/M tokens nearly doubled while accuracy stayed flat.

Why the revert

v0.12.15 vs v0.12.14 per-task cost telemetry on WebBench_READ_v5 (~88 of 198 tasks sampled per run; the 0.12.8-14 lineage tracked tightly in this band):

Metric v0.12.14 v0.12.15 Δ
Success 137/198 (69%) 134/198 (68%) flat (noise)
Cost/M tokens $0.234 $0.453 +94%
Cache-hit rate 62.7% 19.7% -43pp
Cost/task $0.0557 $0.0769 +38%

Mechanism: shrinking the cached prefix invalidated Gemini's implicit cross-task cache. Cache-rebuild cost exceeded the per-call schema savings. Same compensation pattern surfaced in the v0.11.x cycle — any change inside the cached prefix forces a cross-task cache rebuild whose cost exceeds the byte savings.

What's in this release

  • Pure git revert of merge 70a10cb (PR #22).
  • Version bumped 0.12.14 → 0.12.16 across Cargo.toml, Cargo.lock, pyproject.toml.
  • docs/v0.12-plan.md updated:
    • v0.12.15 attempt + result + REJECTED reason.
    • v0.12.16 = pure revert (no code lever; substrate sanity-check).
    • Next slices pivoted away from cached-prefix trims toward per-step uncached bytes.

Validation

  • PYTHONPATH=python pytest tests/test_agent_compat.py tests/test_final_answer_guards.py tests/test_blocked_search_guards.py -q — all 99 tests pass.

Expected eval signature

bu-rust 0.12.16 + WebBench_READ_v5 should reproduce v0.12.14's profile (~69% accuracy, ~$0.234/M tokens, ~63% cache hit). A match confirms the v0.12.15 mechanism and clears the substrate for the next experiment.

v0.12.15

Choose a tag to compare

@sauravpanda sauravpanda released this 23 May 19:31
70a10cb

Summary

v0.12.15 reduces repeated Gemini prompt overhead by hiding duplicate upstream compatibility alias tool declarations from the default LLM schema.

Aliases such as input_text and press_keys are still accepted by runtime dispatch for compatibility, but the model now sees the canonical tool names by default. Callers that need the old full schema can pass Agent(..., expose_tool_aliases=True).

Cost focus

Local schema measurement dropped the default advertised tool declarations from 62 tools / 33,114 bytes to 44 tools / 22,456 bytes, a 32% schema-byte reduction. This targets cached-prefix cost and first-call fresh prompt size while avoiding canonical tool removal in this release.

Validation

  • .venv/bin/python -m unittest tests.test_agent_compat tests.test_prompt_metrics
  • .venv/bin/python -m unittest discover -s tests
  • .venv/bin/python -m compileall -q python tests bench
  • cargo check -p bu-py
  • cargo test -p bu-cdp -p bu-dom -p bu-browser
  • .venv/bin/python bench/release_preflight.py
  • git diff --check

v0.12.14

Choose a tag to compare

@sauravpanda sauravpanda released this 22 May 18:17
727775c

Summary

  • Restores self-validation by default, but only for high-risk final answers.
  • Adds self_validate_risk_only so callers can force validate-every-final behavior when needed.
  • Shortens the validation prompt to reduce the cost of validation turns.
  • Updates the v0.12 plan with the 0.12.13 eval result and 0.12.14 adjustment.

Motivation

0.12.13 removed validation turns and reduced cost/steps, but success fell from 140/198 to 132/198 and Incorrect Result failures increased from 16 to 24. Validation was expensive, but it was catching real bad finals. This release restores validation selectively for search/filter/sort/locator tasks, counted top/first/list tasks, recency/current/latest tasks, price/date/availability/fare tasks, and blocked/external-evidence finals.

Dry-run over the 0.12.13 traces estimates validation would apply to about 73 successful step>=3 tasks, versus 102 provisional finalization turns in 0.12.12.

Validation

  • .venv/bin/python -m unittest discover -s tests
  • .venv/bin/python -m compileall -q python tests bench
  • cargo check -p bu-py
  • .venv/bin/maturin develop
  • cargo test -p bu-cdp -p bu-dom -p bu-browser
  • .venv/bin/python bench/release_preflight.py
  • Installed package reports browser_use_rs.__version__ == "0.12.14"

v0.12.13

Choose a tag to compare

@sauravpanda sauravpanda released this 22 May 17:27
f7facfb

Summary

  • Disables global self-validation by default to remove routine done -> validation -> done confirmation turns.
  • Keeps explicit self_validate=True support for callers that want the old validation behavior.
  • Keeps targeted final-answer recovery nudges and top-N/list count checks enabled.
  • Updates the v0.12 plan with the 0.12.12 eval result and 0.12.13 adjustment.

Motivation

The 0.12.12 run recovered part of the blocked-site success regression, but trace detail showed 100/198 tasks called done twice and 102/198 tasks had a provisional finalization turn. Those provisional-final tasks averaged about $0.0614 versus $0.0487 for the rest. This release removes that global confirmation round trip while preserving targeted guards.

Validation

  • .venv/bin/python -m unittest discover -s tests
  • .venv/bin/python -m compileall -q python tests bench
  • cargo check -p bu-py
  • .venv/bin/maturin develop
  • cargo test -p bu-cdp -p bu-dom -p bu-browser
  • .venv/bin/python bench/release_preflight.py
  • Installed package reports browser_use_rs.__version__ == "0.12.13"

v0.12.12

Choose a tag to compare

@sauravpanda sauravpanda released this 21 May 17:51
ae1b08e

Summary

  • Corrects the 0.12.11 blocked-site fallback regression before the L2 eval.
  • Allows bounded external-result completion for static public facts when target-site variants are blocked and the visible result directly answers the task.
  • Keeps snippet-only answers rejected for live/current data, prices, availability, bookings, account-gated pages, locators, and site actions.
  • Restores softer blocked/search loop force-final thresholds closer to 0.12.10.
  • Updates the v0.12 plan with the 0.12.11 eval result and 0.12.12 adjustment.

Validation

  • .venv/bin/python -m unittest discover -s tests
  • .venv/bin/python -m compileall -q python tests bench
  • cargo check -p bu-py
  • .venv/bin/maturin develop
  • cargo test -p bu-cdp -p bu-dom -p bu-browser
  • .venv/bin/python bench/release_preflight.py
  • Installed package reports browser_use_rs.__version__ == "0.12.12"

v0.12.11

Choose a tag to compare

@sauravpanda sauravpanda released this 21 May 00:35
83f4796

Summary

  • Add a shared blocked-site recovery policy across prompt variants.
  • Budget blocked-site recovery to one targeted search fallback, one same-site fallback URL, and one same-site read/extract attempt.
  • Remove the validation-path instruction that could force a fresh external search while finalizing.
  • Tighten existing bot-block/search-fallback loop windows so blocked/search-engine tails stop earlier.
  • Mark search-engine challenge responses as consuming the search fallback budget.

Validation

  • .venv/bin/maturin develop
  • .venv/bin/python -m unittest tests.test_blocked_search_guards tests.test_final_answer_guards tests.test_batch_guard_handling
  • .venv/bin/python -m unittest discover -s tests
  • .venv/bin/python -m compileall -q python tests bench
  • cargo check -p bu-py
  • cargo test -p bu-cdp -p bu-dom -p bu-browser
  • .venv/bin/python bench/release_preflight.py
  • version import check for 0.12.11

v0.12.10

Choose a tag to compare

@sauravpanda sauravpanda released this 20 May 23:03
98e3f2b

Changes

  • Reuse the previous full PAGE_STATE prompt when URL + cleaned LLM-facing DOM are unchanged.
  • Append a small PAGE_STATE_UNCHANGED marker instead of resending another screenshot/DOM blob, allowing provider prompt caching to reuse the retained full state.
  • Force a full refresh after 3 reused steps by default.
  • Add state_cache_max_reuse_steps, BROWSER_USE_RS_STATE_CACHE_MAX_REUSE_STEPS, and BU_RS_STATE_CACHE_MAX_REUSE_STEPS; set to 0 to disable.
  • Add prompt metadata for page-state bytes, page-state image bytes, reuse marker count, and reuse count.

Notes

  • The cache key uses cleaned browser state: URL + rendered LLM DOM. It does not use raw DOM, and it ignores exact screenshot bytes to avoid animation/ad churn defeating cache reuse.
  • Full PAGE_STATE is not reused when the retained full state contained transient <read_state>.

Validation

  • .venv/bin/maturin develop
  • .venv/bin/python -m unittest discover -s tests
  • .venv/bin/python -m compileall -q python tests bench
  • cargo check -p bu-py
  • cargo test -p bu-cdp -p bu-dom -p bu-browser
  • .venv/bin/python bench/release_preflight.py

v0.12.9

Choose a tag to compare

@sauravpanda sauravpanda released this 19 May 22:50
5f4ca8a

Changes

  • Cap the auto-injected DOM page state at 64 KiB by default to reduce repeated text prompt cost on large pages.
  • Add dom_max_bytes, BROWSER_USE_RS_DOM_MAX_BYTES, and BU_RS_DOM_MAX_BYTES controls; set to 0 to disable the cap.
  • Filter valid action indices to the DOM rows actually shown to the model, so omitted [N] values cannot be used accidentally.
  • Add DOM cap metadata for follow-up evaluation analysis.

Validation

  • .venv/bin/maturin develop
  • .venv/bin/python -m unittest discover -s tests
  • .venv/bin/python -m compileall -q python tests bench
  • cargo check -p bu-py
  • cargo test -p bu-cdp -p bu-dom -p bu-browser
  • .venv/bin/python bench/release_preflight.py

v0.12.8

Choose a tag to compare

@sauravpanda sauravpanda released this 19 May 19:17
931cc24

Summary

  • Reduce Gemini screenshot/media token cost by defaulting ChatGoogle media inputs to low resolution.
  • Honor Agent(..., vision_detail_level=...) by forwarding it to providers with media-resolution support.
  • Honor images_per_step=0 by omitting screenshot image parts from LLM prompts while keeping DOM state.
  • Add modality token details to usage logs and ChatInvokeUsage.model_dump() for image/text cost attribution.

Verification

  • .venv/bin/python -m unittest discover -s tests -q
  • .venv/bin/python -m compileall -q python/browser_use_rs tests bench
  • cargo check -p bu-py
  • cargo test -p bu-cdp -p bu-dom -p bu-browser
  • BROWSER_USE_RS_DISABLE_DOTENV=1 .venv/bin/python bench/release_preflight.py
  • git diff --check

v0.12.7

Choose a tag to compare

@sauravpanda sauravpanda released this 13 May 23:06
  • Restore conservative default web_search behavior after the 0.12.6 eval regression: search opens the SERP and returns immediately by default.\n- Keep the richer SERP snippet extraction available behind BROWSER_USE_RS_WEB_SEARCH_SNIPPETS=1 for isolated testing.\n- Retain the Anthropic Haiku fix so effort is only sent when thinking is active.