Skip to content

v0.12.14

Choose a tag to compare

@sauravpanda sauravpanda released this 22 May 18:17
· 5 commits to main since this release
727775c

Summary

  • Restores self-validation by default, but only for high-risk final answers.
  • Adds self_validate_risk_only so callers can force validate-every-final behavior when needed.
  • Shortens the validation prompt to reduce the cost of validation turns.
  • Updates the v0.12 plan with the 0.12.13 eval result and 0.12.14 adjustment.

Motivation

0.12.13 removed validation turns and reduced cost/steps, but success fell from 140/198 to 132/198 and Incorrect Result failures increased from 16 to 24. Validation was expensive, but it was catching real bad finals. This release restores validation selectively for search/filter/sort/locator tasks, counted top/first/list tasks, recency/current/latest tasks, price/date/availability/fare tasks, and blocked/external-evidence finals.

Dry-run over the 0.12.13 traces estimates validation would apply to about 73 successful step>=3 tasks, versus 102 provisional finalization turns in 0.12.12.

Validation

  • .venv/bin/python -m unittest discover -s tests
  • .venv/bin/python -m compileall -q python tests bench
  • cargo check -p bu-py
  • .venv/bin/maturin develop
  • cargo test -p bu-cdp -p bu-dom -p bu-browser
  • .venv/bin/python bench/release_preflight.py
  • Installed package reports browser_use_rs.__version__ == "0.12.14"