Skip to content

Benchmarks

Steven Enamakel edited this page Sep 1, 2026 · 6 revisions

Benchmarks

What a room of agents buys over one agent answering alone, what it costs, and which settings decide the difference.

Everything here is produced by:

cargo run --release -p tinyhivemind-hive --example bench

The harness itself is documented next to the code, in crates/tinyhivemind-hive/examples/bench/. Episode policy is the tuning guide these numbers produced, and Further reading has the literature they sit against.

Summary

Five agents, four options, 5000 seeded rooms, one core:

arm       turns/ep   decided %   correct %       ns/step    episodes/s
ladder        1.00       100.0        57.6          1109        901660
vote         15.00       100.0        78.5             0           inf
hive          6.16        89.7        73.3          2231         62637
hive+         6.75        99.4        82.1          2278         56641

A tuned deliberation is right 82.1% of the time in 6.75 turns. The matched-budget control, given all fifteen turns, reaches 78.5%. One responder off the ladder, which is how the system behaves today, reaches 57.6%. The state machine costs about 2.3 microseconds per step, six orders of magnitude below a model turn.

The margin over the control is a few points, and it is meant to be. Independent sampling plus a plurality is most of what a room is for, and a protocol that could not clear that bar would not be worth its budget.

What is measured

A room chooses between several options, exactly one of which is genuinely best. Every member holds a private, noisy evaluation of every option: the true quality plus a uniform error of half-width --noise. No member is individually reliable, so the room's only route to the right answer is to pool what its members separately believe.

The participants are arithmetic, deliberately. A language model would make the numbers unreproducible and would confound protocol quality with model quality. What is being measured is whether a policy aggregates information or throws it away, which is the question a host has to answer when it configures a desk. Real models appear further down, where the claim is only that they can hold the protocol.

Everything is seeded. The same --seed produces the same rooms, the same private evaluations, and the same transcripts, so a change in a reported number is a change in the library rather than in the weather.

The arms

Every arm decides the same rooms from the same private evaluations.

arm what it is turns
ladder responder_plan selects one responder off the real ladder, selector rung included and validated through accept_selection, and that agent answers alone 1
vote independent answers decided by plurality, nobody seeing anybody: self-consistency at a matched budget the whole budget
hive a deliberation episode at EpisodePolicy::DEFAULT up to the budget
hive+ the same, at the tuned policy up to the budget
hive+ref the tuned policy with refutation_cap: Some(2) up to the budget
hive+ev the same, plus require_evidential up to the budget

vote is the honest control, and Condorcet's jury theorem is the reason it is a strong one. A multi-agent result without such a control is close to meaningless, because the multi-agent arm has usually just spent more compute. It is given the whole budget, which is more turns than the deliberation actually spends, though with deterministic participants it saturates at one distinct answer per member, which is exactly what self-consistency does with a deterministic sampler.

Correctness is scored over the whole sample, including episodes that decided nothing. An arm cannot buy accuracy by declining to answer.

Results

The two bounds on quorum

The quorum threshold is the single most consequential setting, and it has a bound on each side. Five members, 5000 rooms, everything else held at the tuned policy:

quorum deadlocked exhausted decided % correct %
2 of 5, below a majority 514 0 89.7 73.3
3 of 5, smallest majority 0 29 99.4 82.1
5 of 5, unanimity 0 2092 58.2 55.2

Below a majority, rooms deadlock. Five members can put two grounded supporters behind each of two options, and an episode in which two options both carry is deadlocked by definition: no amount of further support resolves it, because both stay above the line. Requiring a majority makes that state unreachable and the deadlock rate falls to zero.

At unanimity, rooms cannot finish. Cross-inhibition removes a silenced advocate from a topic's supporter set and does not put them back, so a single grounded !object makes quorum unreachable for the rest of the episode. Two in five episodes then spend their whole budget without deciding. A live three-member room hit exactly this, described below.

So: a majority of the desk, and never the whole of it. The benchmark's tuned policy computes threshold = min(agents / 2 + 1, agents - 1).

Refutation and evidential grounds lose

Both are opt-in and both are off in QuorumPolicy::DEFAULT, and this table is why. 5000 rooms, five members, --noise 90, everything else at the tuned policy:

arm turns/ep decided % correct %
vote 15.00 100.0 78.5
hive+ 6.75 99.4 82.1
hive+ref 8.99 88.6 75.0
hive+ev 10.29 60.8 55.9

!refute costs seven points and drops the arm below even the vote control. require_evidential costs twenty-six and fails to decide two episodes in five, because a room that has not deposited facts cannot carry anything at all.

The damage scales with how noisy each member's private evaluation is — nothing at ±30, six points at ±60, seven at ±90, fifteen at ±120 — and that is the diagnosis. A refutation is global where an objection is local. An !object removes one advocate from one topic; a !refute caps the topic for the whole room, so a member firing one on a noisy read removes an option for everybody. This is the same neutrality the live rooms showed when their one observed !object fired against the correct option, with a much larger blast radius.

The grid search agrees without being asked: --sweep scores 864 policies over the same rooms, and every policy in the top twelve has both knobs off.

What this does not test is the case the mechanism was built for. The simulated task gives every member a noisy estimate of every option, so there is no decoy that accumulates support no individual's private read contradicts and no fact held by one member that overturns it — which is what a hidden profile is, and what the live checkout-503 scenario has. On a task where every member can already evaluate every option, weighing evidence against support has nothing to win and a real cost to pay. The full record, including the open items, is in docs/experiments/2026-09-01-refutation-and-grounds.md.

The budget has to scale with the desk

A fixed budget makes a larger room look worse than a smaller one, and the effect is entirely an artifact of the cap. An eight-member room, 1500 rooms each:

budget decided % correct % turns actually spent
12 65.3 63.1 10.36
16 89.6 82.9 10.96
20 94.5 86.3 11.24
24 96.4 87.7 11.42

A blind opening round costs one turn per member before anybody has seen anybody, a majority then has to assemble on one option, and the decision has to be recorded. Three turns per member covers that, and it is a cap rather than a cost: the eight-member room finishes in 11.4 turns of the 24 it is allowed. At five members, budgets of 15, 20 and 25 score 82.0, 82.1 and 82.1. Past the point where the room can finish, extra budget buys nothing.

The blind round is not decoration

Turning it off, five members, 5000 rooms:

opening round decided % correct %
blind 99.4 82.1
full visibility 100.0 58.0

With full visibility from the first turn the room cascades onto whatever was proposed first and lands level with a single agent. That is an information cascade, and Visibility::Blind is what prevents it, bought as a filter on the projection rather than as concurrency. See ADR 0002.

Across desk sizes

2000 rooms per size, threshold and budget scaled as above:

agents quorum budget ladder % vote % hive+ % turns/ep decided %
3 2 9 57.3 68.4 71.6 4.42 99.8
4 3 12 57.3 74.2 76.8 6.68 94.6
5 3 15 57.6 78.8 81.5 6.79 99.3
6 4 18 57.1 82.2 83.4 9.04 95.6
8 5 24 58.4 87.3 88.6 11.32 96.9

The deliberation beats the matched-budget control at every size, by 1.2 to 3.2 points, while spending roughly half the turns. Deadlocks are zero throughout.

What the library costs

ns/step is one call to tinyhivemind_hive::step over a live transcript, with the participants' own time excluded. That is about 2.3 microseconds, or roughly 57,000 whole episodes per second on one core. An episode of nine steps costs about 20 microseconds of library time. A model turn is six orders of magnitude more expensive, so the protocol is free in any real deployment.

Three changes made during this work cut that cost by about a fifth, measured before and after under identical settings (2816 to 2222 ns/step). None of them changes behaviour, and every arm's outcome was byte-identical across them:

episode::step folds a borrowed Vec<&SessionMessage> rather than cloning the filtered transcript on every step, and computes consensus once instead of twice. trace::extract returns early on a body containing no !, which is most of a real transcript, before scanning for fences. quorum::standings folds on borrowed topic and agent keys and allocates owned strings only for what survives, and attention::bids hoists saturation and the reader-independent half of salience out of its per-member loop.

The remaining cost is dominated by re-reading the transcript on every step, which is inherent. An episode is a pure fold with no state cached between calls, so a step over a transcript of n messages parses n messages.

Live agents

--agent-cmd swaps the simulated participants for a real agent CLI, one process per authorized turn. These runs used opencode 1.18.25 against OpenRouter:

cargo run --release -p tinyhivemind-hive --example bench -- \
  --agent-cmd "opencode run --pure -m openrouter/~openai/gpt-mini-latest" --agents 5

They assert nothing. The deterministic arms measure the protocol, and a handful of live episodes could not measure anything. What they establish is that real models hold the trace grammar, and what they surfaced is a set of host-side obligations the simulation cannot see. All four were fixed in the harness, and each fix is one a host owes its agents rather than something the library can impose:

what happened live the fix
Models coined #rollout and #rollout-strategy for one idea; support split across two names never adds up to a quorum. The prompt names the options already on the floor, folded through the library's own standings, with how many supporters each holds and how many it needs.
Four consecutive turns restated the same !question verbatim. repetition_cap damps a restated support and cannot see this. The prompt shows a participant its own last line and asks for something that moves the room on.
Models wrote !commit while the room was still deliberating, which adds no supporter; one episode spent its whole budget recording a decision it never reached. The prompt offers only the moves that count in the turn's phase, so !commit appears only under Phase::Commit.
Four of five proposals dropped the #, so the lines named no topic and deposited nothing. The grammar's sigils are called out explicitly, with a right and a wrong example.

A three-member room also exhausted its budget before the policy fix above, because a threshold of three on a three-member desk is unanimity and one !object had silenced an advocate. That is the second bound on quorum, found live before it was measured.

After the fixes, a five-member room converged on #rollout in 6 turns and 18 seconds of wall clock, with every proposal well formed and one commit. A three-member room converged in 5 turns and 14 seconds.

A real problem, and a control that can lose

Everything above measures whether a model can hold the grammar. None of it measures the thing the protocol exists for, because the synthetic brief has no answer and every member is told the same nothing. --scenario adds a problem that does:

cargo run --release -p tinyhivemind-hive --example bench -- \
  --agent-cmd "opencode run --pure -m openrouter/openai/gpt-5-mini" \
  --scenario crates/tinyhivemind-hive/examples/bench/scenarios/checkout-503.txt \
  --repeat 5

A scenario file carries a shared brief, the options under the ids the room should use, a private brief per member, and the recorded answer. The private briefs never reach the shared journal. A fact every member can already read is not private information, and a room whose members all start from the same facts has nothing to pool — which is the quiet reason most multi-agent results are uninteresting. Both arms then run against the same real agents: a deliberation episode, and an independent poll of the same members answering alone.

The control has to be able to lose

Two scenario designs were thrown away before one separated the arms, and how they failed is the most portable finding here.

Both made the correct option #retries — a client retry storm after a release. That is the canonical 503 story, and a language model reaches for it unprompted. Every member solved it alone and the vote scored five out of five in both designs. An arm that cannot lose is not a control, and a scenario whose answer survives deleting every private brief is not measuring deliberation at all.

The third design is a hidden profile: the brief itself plants the retry storm, and the recorded answer is the boring one, #pool. Reaching it needs four facts held by four different members — the pool caps at 20 and a request waiting 250ms is failed with a 503; each in-flight request holds one connection for its whole life; in-flight requests have sat at 24–31 since 09:14 with traffic flat; the release quintupled how long each request lasts — plus a fifth that kills the decoy, that the retry path shipped disabled and has never fired. No member holds two of the four, and any one alone is inert.

What the rooms did

Five members, quorum 3, budget 15, one episode per row:

agent CLI hive turns vote plurality
claude -p --model sonnet #pool, correct 10 #retries, wrong (3/5)
opencodeopenai/gpt-5-mini #pool, correct 8 #retries, wrong (4/5)
opencodeopenai/gpt-5-mini #pool, correct 8 #retries, wrong (3/5)
opencodeqwen3-235b-a22b #pool, correct 11 tied 2–2–1, no answer

Four rooms reached an answer no member could reach alone, at 8–11 turns of a 15-turn budget and 45–75 seconds of wall clock. The poll never once returned the recorded answer.

Ten further episodes on one model, gpt-5-mini, put a rate on that:

arm            hive correct   turns/ep   vote correct
deliberation        6/10          7.6          —
independent poll      —            —          0/10

The poll lost all ten, always to #retries. Across every episode ever run on this scenario it has returned the recorded answer zero times out of sixteen, which is the point of a hidden profile rather than a surprise: a member voting alone has its own facts and the brief, and both point at the decoy.

Deliberation is not reliable here either. Six in ten is what a room of this model achieves at this budget, and the four failures all fail the same way, described below. The honest summary is that the shape separates the arms cleanly and reproducibly — not that deliberation solves it.

Support is counted; grounds are not weighed

This is the finding worth acting on, and it comes from the rooms that got it wrong.

In all four failed episodes of the ten the decisive refutation — the retry path shipped disabled, zero retries fired today — was already in the transcript, at a sequence every later message could cite. It changed nothing. Members went on depositing !support #retries ^6, and three grounded supporters carried the decoy. A room fails this task in exactly one way, and it is this one.

The library counts distinct grounded supporters. It does not, and as a pure fold cannot, check that a cited message supports the claim it is cited for. So far so correct. But the grammar gives a member no way to say the thing that would have mattered: this evidence refutes that topic. !object >N ^M silences one advocate message. Killing a hypothesis with a fact means objecting to every advocate separately, one turn each, and a room runs out of budget before it runs out of advocates.

A negative evidence-to-topic link — a !refute #topic ^N that debits or caps a topic's standing rather than silencing one author — is the missing move. It is a fold over the same traces and needs no port. Whether it should gate quorum or only weight it is a design question and wants an ADR before an implementation.

Two more things the live rooms showed

A member's fact can attach to the wrong hypothesis. The auditor's concurrency evidence is compatible with both the true cause and the decoy. In the failed rooms the auditor deposited it and proposed #retries from it, and every later member cited that message as grounds for the decoy. Pooled information does not help if it lands under the wrong heading, and nothing in the protocol notices.

Cross-inhibition fires, and it can fire against the truth. The one live !object observed across every run was aimed at #pool — the correct option — arguing that raising the pool masks a duration problem. The mechanism worked exactly as specified and removed a supporter from the right answer. It is not wrong; it is neutral, and a benchmark that reports only its wins is not reporting it.

Two harness defects the runs exposed

gpt-5-mini copied the prompt's <one sentence> placeholders into its output verbatim, angle brackets and all, where they survived into every later citation of that message. claude -p produced the other one: !support #retries ^5 … points at #pool instead, a support marker on one topic whose prose argues for another. The library counted a supporter the author plainly did not intend.

Both prompt defects were repaired, and the repair was then run against the prompt it replaced: five episodes each, same model, same scenario. The arms tied at three correct out of five, 7.8 turns against 7.4. The change is therefore not detectably better or worse, and it is kept for the transcripts it produces rather than for any accuracy claim. An earlier four-correct-then-two- wrong split had suggested a regression; at ten episodes it was noise.

The vote control had a defect of its own. The qwen room polled two, two and one, and the harness reported a #pool plurality because the tally sorted stably and #pool was inserted first. A tie is not a decision, and resolving it by arrival order hands the control a win it did not earn. plurality now returns nothing on a tie.

A federation of desks

--swarm measures something the arms above cannot: several desks that cannot read each other's transcripts, deciding one question. See Cross-desk referral for the mechanism.

The task changes shape to make the boundary cost something. Each desk overrates one option — a different one per desk — because its members read the same transcript and are wrong about the same thing. Within a desk that bias is invisible and averaging correlated error does not remove it; across desks the biases cancel. Three desks of four, 400 seeded federations:

arm        correct  decided     turns  crossings
siloed        0.2%        1      15.9        0.0
swarm        77.5%      389      32.3       12.0
pooled       74.5%      371      16.7        0.0
merged       10.5%       96      33.8          —
vote          4.0%      141      12.0          —

siloed is the same desks with referrals off, and it is not merely worse — it is destroyed. 1,199 of its 1,200 desk episodes reach a confident decision, and three confident desks disagreeing three ways produce no plurality at all. A federation of well-run rooms that cannot talk does not degrade gracefully.

merged puts all twelve members on one desk with the whole budget, and scores 10.5%: a larger room with three factions cannot assemble a majority quorum, so most episodes exhaust. Removing the boundary is not the fix, and costs the same turns as crossing it. pooled is the ceiling control — every desk handed every other's readings for free — and swarm matches it, so the protocol delivers essentially all of what the information is worth and what it costs is turns. At --bias 0, where no desk has a blind spot, every arm scores 100% and crossing buys nothing at twice the turns; that is why every knob in ReferralPolicy::DEFAULT is off.

The largest effect measured is not in the library. A desk whose members share a bias reaches quorum inside its own blind opening round, so a fact arriving after that is one the desk has already voted past: asking before backing anything rather than after is the difference between failing outright and 77.5%.

Five live runs against claude -p --model sonnet add what the simulation cannot. The mechanism works end to end with real agents, and in the best run a desk reached an answer no member of it could have reached alone by asking another desk for a number. But agents ignored the move entirely until it was placed in the marker list rather than above it, and one desk answered a question with its hypothesis instead of its evidence and exported the error intact. A protocol that moves messages does not by itself move evidence. Full write-up, sweeps and transcripts: docs/experiments/2026-09-02-federated-hidden-profile.md.

Reproducing

cargo run --release -p tinyhivemind-hive --example bench                      # the table above
cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000
cargo run --release -p tinyhivemind-hive --example bench -- --agents 8
cargo run --release -p tinyhivemind-hive --example bench -- --quorum 5        # unanimity
cargo run --release -p tinyhivemind-hive --example bench -- --no-blind        # the cascade
cargo run --release -p tinyhivemind-hive --example bench -- --sweep           # the policy grid
cargo run --release -p tinyhivemind-hive --example bench -- --trace           # one episode
cargo run --release -p tinyhivemind-hive --example bench -- --swarm           # a federation
cargo run --release -p tinyhivemind-hive --example bench -- --swarm --bias 0  # no blind spots
cargo run --release -p tinyhivemind-hive --example bench -- \
  --agent-cmd "opencode run --pure -m openrouter/openai/gpt-5-mini" \
  --scenario crates/tinyhivemind-hive/examples/bench/scenarios/checkout-503.txt \
  --repeat 5                                                                  # a real problem

CI runs cargo run -p tinyhivemind-hive --example bench -- --episodes 25, so the harness cannot rot. Live mode is not in CI and needs a configured agent CLI.

What this does not show

Nothing about model quality. The deterministic participants are arithmetic. A room of language models may aggregate better or worse than this, and these numbers cannot tell you which.

Nothing about real tasks, in the deterministic arms. One synthetic task with a known best option and independent errors is the friendliest possible case for aggregation. Real disagreements are correlated, and correlated errors are exactly what pooling cannot fix — within the group that shares them. The federated arms above are the one place that limitation is measured rather than assumed, and they say what to do about it: pool across a boundary the correlation does not cross. They also say what not to expect, because two desks sharing a blind spot would confirm each other rather than correct each other, and that case is not measured anywhere here.

The scenario runs above are a first step off that synthetic task and are not a substitute for it. Four live rooms are an existence proof that a hidden profile separates deliberation from a matched-budget poll; they are not a rate, and the two rooms that failed did so by pooling their information under the wrong hypothesis, which is a failure mode the deterministic arms cannot produce at all.

Nothing about long rooms. Conformity in a group of language models rises with interaction time. The budgets here are small on purpose, and a longer episode should be expected to buy correlated error rather than better judgement.

tinyhivemind-hive is a protocol for bounded deliberation with an auditable termination reason. That is the whole claim.

Clone this wiki locally