Skip to content

BU Bench V2.1.1 — current 200-task runner

Latest

Choose a tag to compare

@MagMueller MagMueller released this 25 Sep 05:42
af6c7f7

BU Bench V2.1.1 includes all merged updates since v2.1: the default BrowserCode runner, Cloud and local browsers, Luna xhigh, 100-task concurrency, CAPTCHA-aligned rubrics, the dataset cleanup, and the validated Reuters playback redesign.

Download Source code for the complete runnable repository, or follow the setup instructions:

git clone --branch v2.1.1 --single-branch https://github.com/browser-use/benchmark.git
cd benchmark

After installing the documented dependencies and setting your own API keys, run:

uv run python run_eval.py

The default is all 200 tasks, BrowserCode 0.1.20, OpenAI GPT-6 Luna xhigh, Browser Use Cloud, the GPT-5.6 Luna xhigh findings judge, 3,600 seconds per task and up to 100 concurrent tasks. Local Chrome and optional Laminar reporting remain supported. GitHub Actions runs the same entry point.

Dataset and scoring

  • Includes the CAPTCHA/source-verification corrections and the Reuters task 185 playback redesign. All 200 task IDs remain; only task 185 changes its item layout and weights.
  • The dataset and encrypted review cases are byte-identical to the reviewed, merged Reuters revision. No new evaluation or regrade was performed for packaging.
  • Historical scores are unchanged. One real task 185 run scored 100/100 with 25.54 seconds of verified playback and a genuine final screenshot. This is a single successful execution, not a reliability estimate.
  • Other saved-evidence semantic reviews remain pending, as listed in the revision record.

Downloads

The attached BU_Bench_V2.enc is the current 200-task dataset. BU_Bench_V2_review_cases.enc contains review scenarios, not runnable tasks. rubric_revision.json records release and content provenance; SHA256SUMS verifies the attached files. The old 55-task subset is not included.

Content revision: 2026-09-25-reuters-playback.
Dataset SHA-256: 0c014b056192e261b06f04ed6da6dcf37f2135150cf58c87ea839bd4e150e5c3.

The original v2.1 tag and assets remain unchanged for older comparisons. Keep decrypted tasks, rubrics and raw evidence private.