Skip to content

docs: refresh BU Bench harness plot with current model results - #143

Merged
Alezander9 merged 1 commit into
mainfrom
refresh-bu-bench-harness-plot
Aug 4, 2026
Merged

docs: refresh BU Bench harness plot with current model results#143
Alezander9 merged 1 commit into
mainfrom
refresh-bu-bench-harness-plot

Conversation

@Alezander9

@Alezander9 Alezander9 commented Aug 4, 2026

Copy link
Copy Markdown
Member

What

Copies the current browser_harness_by_model_{light,dark}.png from
benchmark-x-laminar@45b563b into static/.

The versions in the README were generated 2026-07-29 and are several
generations behind. Changes now reflected:

  • Added: claude-opus-5 (87.0%), claude-fable-5 (87.0%), grok-4.5 (86.3%), kimi-k3 (86.0%), gpt-5.6-sol (84.0%), gpt-5.6-luna xhigh (82.0%) / default (72.4%), gpt-5.6-terra (72.0%), qwen3.8-max (79.0%), deepseek-v4-flash-0731 (76.0%)
  • Dropped: base deepseek-v4-flash (superseded by -0731), gpt-5.6-luna-medium, and the stale mimo/glm-4.7/kimi-k2.6/qwen3.6-plus/grok-4.3 bars
  • Open-weight bars changed from a solid blue fill to a gray fill with a blue dashed outline

Not in this PR

Two things above the bar chart are also stale and need a decision before
they can be propagated:

  1. browsercode_best_models_{light,dark}.png (the radial infographic) is
    hand-authored HTML upstream, not generated from the results, so it cannot be
    mechanically refreshed. It still shows Opus 4.7 / GPT-5.5 / GLM 5.2 /
    Gemini 3.1 Pro / Minimax M3, and its numbers no longer match the results
    (e.g. Minimax M3 78.0% vs 74.0% measured).
  2. The "Recommended models" bullet list mirrors that infographic's five
    categories. Against current data, three of the five have changed:
    best speed gpt-5.5 (17.4 tasks/hr) -> gpt-5.6-terra (29.3);
    best open-weight glm-5.2 (84.0%) -> kimi-k3 (86.0%);
    lowest cost minimax-m3 -> deepseek-v4-flash-0731 ($1.60/100 tasks).
    Left unchanged here so the list does not contradict the infographic
    directly above it.

Summary by cubic

Refreshed the BU Bench harness chart images to match the latest upstream results, so the README shows current model performance. Adds the newest model bars, removes outdated ones, and updates open‑weight bar styling (gray fill with a blue dashed outline).

Written for commit 6f40bce. Summary will update on new commits.

Review in cubic

Propagates browser-use/benchmark-x-laminar@45b563b. The README copies were
generated 2026-07-29 and predate seven changes to the plot:

  - claude-opus-5 (87.0%), claude-fable-5 (87.0%), grok-4.5 (86.3%),
    kimi-k3 (86.0%), gpt-5.6-sol/terra/luna, qwen3.8-max (79.0%) and
    deepseek-v4-flash-0731 (76.0%) added
  - base deepseek-v4-flash and gpt-5.6-luna-medium dropped
  - open-weight bars are now a gray fill with a blue dashed outline
    instead of a solid blue fill

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 2 files

Re-trigger cubic

@Alezander9
Alezander9 merged commit 4db4b32 into main Aug 4, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant