Skip to content
Drew T edited this page Sep 29, 2026 · 3 revisions

The bench

/bench measures the agents themselves. It runs fixed fixture tasks through each agent headlessly and grades them without a model in the loop: hidden tests decide the coder tasks, including a bug-hunt tier on a 2,000-line fixture repository; a pinned fixture coder carries out an expert's brief and the hidden tests decide whether the brief was right; planners, the critic, review and the router are checked against their contracts, and the planners, the critic and review also get a blind rubric score from a pinned grader. Suites: light (the default), full, canary, or one agent. Every result records the agent's model, effort and body hash, the fixture hash, the harness version and your plan's preset.

The public series

The public series is in the repository's bench/results/. The project's own Action runs the light suite every week and the full suite every fourth week, and publishes each result there.

Running it yourself

/bench reads that series first and compares each installed agent with it: match, update, run-offered, lightly-altered or no-series. It suggests a run only for an agent the series cannot speak for, one you edited or wrote yourself, and runs only when you say so, within a budget of your five-hour window: the light suite stays under 4 % of it and the full suite under 10 %. Your own runs stay on your machine.

Benchmarking a new model

When a new model comes out, PA3 decides whether to use it with the same bench, in five steps:

  1. Price it. Its list price per million tokens goes into the price table, beside the models the agents already run, along with how your plan's usage limits count it.
  2. Make one arm per candidate: a copy of the installed agent with only its name, model and effort changed, so the model is the only thing that differs. The arms never replace the installed agents.
  3. Run every arm on the same fixtures, each task at least twice, and let bench.py check confirm the budget, the graders, the noise between repeats and the calibration.
  4. Read the result against the ceiling. When every arm passes every task, the suite shows that the candidate does not regress and what it costs, but not which model is better; that takes a harder, continuously scored tier.
  5. Promote only on the developer's word: a promotion changes the model a preset gives an agent, and every user of that preset gets it with the next update.

An upgrade is complete per model version. When a new version ships, every agent that runs on its predecessor gets its own A/B (bench.py ab, the full suite, each task twice, the two arms differing only in the model line), and each agent moves on the developer's word, one by one. An agent the bench has no tier for is still listed, named as unbenched, so the result accounts for every agent on the old version. The promotion is done only when a grep of wiki/, the READMEs, INSTALL.md and docs/project-architect.md for the old model name finds it in history alone, such as the dated tables on this page. A switch to another model family, as when retriever-code moved from Haiku 4.5 to Sonnet 5.5, is a separate decision of the developer's and lies outside this rule.

These evaluation runs are not part of the public series, which tracks the agents as installed.

Opus 5.5: what moved to it, and why

Opus 5.5 costs $4 input and $20 output per million tokens, below Opus 4.6's $5 and $25. Three promotions came with it:

  • The coder and the default expert moved to Opus 5.5 at medium effort on 2026-09-23, by the developer's decision; Fable 5.1 stayed as the expert for the hardest tasks. No bench stood behind this one.

  • Both planners moved from Fable 5.1 to Opus 5.5 at medium on 2026-09-24, after a blind planner bench: drafts from each arm, graded against a written rubric (17 points) with the arms hidden until the last sheet, four drafts per arm.

    Planner Opus 5.5 medium Fable 5.1 medium
    Phase plans: mean score (spread) 14.25 (sd 0.50) 13.25 (sd 1.71)
    Phase plans: cost and time per draft $1.29, 221 s $2.73, 323 s
    Generation plans: mean score (spread) 15.0 (sd 1.41) 14.75 (sd 0.96)
    Generation plans: cost and time per draft $0.76, 181 s $2.55, 291 s

    The scores were level within the noise, and Opus cost 47 % and 30 % of Fable per draft.

  • /discuss and the phase closer moved to Opus 5.5 on 2026-09-25, by the developer's decision.

Sonnet 5.5: four agents moved, the coder did not

Sonnet 5.5 arrived on 2026-09-28 at $2 input and $10 output per million tokens, half of Opus 5.5, with cache reads at the same $0.20. It was benched as the coder at medium and at high effort against the installed coder, Opus 5.5 at medium.

On the full coder suite (28 tasks, two attempts each) every arm passed every attempt:

Coder Full suite Cost per attempt Median time Five-hour points
Opus 5.5 medium 56/56 $0.127 28.5 s 3
Sonnet 5.5 high 56/56 $0.066 23.0 s 1
Sonnet 5.5 medium 56/56 $0.045 14.4 s 1

The light suite agreed (28/28 for each arm). The installed coder has passed every task of this suite since 2026-09-25, so the suite was at its ceiling: it showed no regression and the saving, not which model codes better. As a control, Haiku 4.5 passed 44 and 45 of 56 attempts in two runs, at $0.16 an attempt: dearer than Sonnet 5.5 and a fifth of its attempts lost.

A harder tier settled it. The max tier is private: it gives the coder one function of a private matching-decompilation project as MIPS assembly, asks for C, and scores the attempt by compiling it and comparing the instructions with the original's (1.00 when every instruction matches). It ran 15 functions of 53 to 246 instructions, two attempts each, and a hard batch of 5 larger functions (284 to 541 instructions), where Haiku did not run:

Coder Max tier: mean score (exact) Cost per attempt Hard batch: mean score (exact) Cost per attempt
Opus 5.5 medium 1.000 (30/30) $0.41 0.984 (7/10) $1.71
Sonnet 5.5 high 0.971 (28/30) $0.33 0.941 (5/10) $1.73
Sonnet 5.5 medium 0.843 (17/30) $0.24 0.820 (3/10) $0.82
Haiku 4.5 0.444 (1/30) $0.93 not run

Why it was not promoted (the developer's decision, 2026-09-29): Sonnet 5.5 at high scores below Opus 5.5 on hard work, 0.971 against 1.000 on the base set and 0.941 against 0.984 on the hardest functions. On those functions it also costs as much per attempt ($1.73 against $1.71): it reads about 60 % more tokens per attempt (2.9 million against 1.8 million), almost all of it context it has already seen, which costs the same $0.20 per million on both models. At medium it drops to 0.843. Every preset keeps Opus 5.5 at medium as its coder for now, and Haiku 4.5 is not a coder.

The code retriever did move. On 2026-09-29 it was benched on the full retriever suite (28 tasks, two attempts each) against itself, the arms differing only in the model line:

Code retriever Attempts Tasks Cost Turns per answer Time per answer
Haiku 4.5 55/56 27/28 $1.34 5.5 16.0 s
Sonnet 5.5 medium 56/56 28/28 $1.37 3.1 9.2 s

Sonnet 5.5 answered every question for about the same money, in fewer turns and in a little over half the time, so retriever-code now runs on Sonnet 5.5 at medium effort (results in .run/rcode-ab-results, not committed). This was a family change the developer decided on its own, outside the per-model rule.

The three agents that ran on Sonnet 5 moved to Sonnet 5.5 the same day: the router (the pa-session), the document retriever and the web retriever. Each was benched on its full suite (two attempts per task) against itself on Sonnet 5, at medium effort, the arms differing only in the model line. Per-token prices of the two models are the same.

Router (pa-session) Attempts Tasks Cost Turns
Sonnet 5 28/30 13/15 $0.88 1.8
Sonnet 5.5 29/30 14/15 $0.77 1.8
Document retriever Attempts Tasks Cost Turns per answer
Sonnet 5 27/28 13/14 $1.54 4.3
Sonnet 5.5 28/28 14/14 $1.44 3.1
Web retriever Attempts Cost Turns per answer
Sonnet 5 0/28 $0.69 1.4
Sonnet 5.5 0/28 $0.61 1.1

The router tied on 14 of 15 tasks and did better on one; the document retriever tied on 27 of 28 attempts and did better on one; both cost less. The web retriever's table measures nothing about quality: the retriever suite asks questions about a codebase, and the web retriever holds only web search, web fetch and write, so every task fails by construction on both models. The run shows only that nothing in the harness broke. It moved with the family, and the public series has never carried a web retriever arm. So all four Sonnet agents (router, code, document and web retrievers) now run on Sonnet 5.5; the coder stays on Opus 5.5.

Clone this wiki locally