-
Notifications
You must be signed in to change notification settings - Fork 4
Benchmarks delegation
The delegation half of the deliberation benchmark: what happens when a room's members are not interchangeable — one of them holds the fact that decides the question, and the others do not.
It answers three questions, written down before any of these numbers existed.
-
Q1 — does the floor reach the member who holds the fact?
fact %is the share of episodes in which the decisive member deposited a topiced!evidenceline before the commit boundary, andto-factis the mean turn index at which it did. A deposit landing at or after the boundary is compute the room paid for and could not use, and is scored as a miss. -
Q2 — does an informative router beat a uniform ladder?
route %, over the two ladder arms. -
Q3 — what does accuracy cost, once seats have prices?
cost/epandcorrect/kU, right answers per thousand units, under--cost-tiers.
A fourth number, rho, is an obligation rather than a question:
the spec
requires this benchmark to publish the rank correlation between a member's
directory weight and its share of the episode's turns, and to report the
mechanism as having failed if it merely tracks who talked.
| flag | what it does |
|---|---|
--specialists N |
N members read one topic each far more tightly than everybody else, and everybody else's read of that topic widens to match. Information is redistributed, not created. |
--hidden-profile |
One decoy is planted above every member's own argmax except one member's, and that member alone holds the fact that rules it out. |
--blind-evidence |
A member's first turn, while the room is blind, is a deposit rather than a position. Off by default. |
--defer-cap N |
Turns a member may spend saying a topic is not theirs. |
--cost-tiers |
A specialist's turn costs ten units against a lay member's one. |
--history N |
Prior episodes of hive+ from which ladder+dir earns its directory. |
hive+dir is the tuned policy with
the directory on, so BidReason::Knows is reachable.
hive+defer adds defer_cap with nothing routing the vacated turn, and
hive+dir+defer is both. ladder+dir is the responder ladder with each
candidate's earned directory lines rendered into its description and the
contested topic named in the request — the routing rule a subagent description
and a role string implement — validated through the real accept_selection.
hive+cost and all-reasoning appear under --cost-tiers only.
5000 rooms, ±90 noise. With homogeneous expertise there is nothing to route on, and the spec predicted exactly that.
ordinary opening --blind-evidence
5000 rooms correct % 95% CI correct % 95% CI knows %
ladder 57.6 56.2–59.0 57.6 56.2–59.0 —
vote 78.5 77.4–79.6 78.5 77.4–79.6 —
hive+ 82.1 81.0–83.1 75.3 74.1–76.5 0.0
hive+dir 82.1 81.0–83.1 75.7 74.5–76.8 75.1
hive+defer 82.1 81.0–83.1 75.3 74.1–76.5 0.0
hive+dir+defer 82.1 81.0–83.1 75.7 74.5–76.8 75.1
ladder+dir 49.5 48.1–50.9 98.8 98.5–99.1 0.0
Under the ordinary opening the delegation arms are hive+ to the digit and
knows % is 0.0: BidReason::Knows never fires. It needs a member that
is the directory's top holder of the contested topic and has taken no
position on it, and where every member opens with a !propose those never hold
together. ladder+dir is eight points worse than the uninformed draw over the
same five candidates.
ordinary opening --blind-evidence
5000 rooms correct % 95% CI correct % 95% CI knows % route %
ladder 52.6 51.2–54.0 52.6 51.2–54.0 — 18.9
vote 71.1 69.9–72.4 71.1 69.9–72.4 — —
hive+ 74.2 72.9–75.4 67.6 66.3–68.9 0.0 —
hive+dir 74.2 72.9–75.4 68.0 66.7–69.2 78.3 —
hive+defer 74.3 73.1–75.5 67.2 65.9–68.5 0.0 —
hive+dir+defer 74.3 73.1–75.5 68.1 66.8–69.4 79.8 —
ladder+dir 45.1 43.8–46.5 84.4 83.4–85.4 0.0 22.3
ladder+dir routes to the decisive member more often than ladder does —
22.3% against 18.9% — and is seven and a half points less accurate. Directory
weight on a topic is earned by grounding it, so the heaviest holder is the
member who argued it hardest rather than the one who reads it best. Routing
precisely to the wrong criterion is worse than not routing.
Under --blind-evidence the specialist deposits before the commit boundary in
93.4% of episodes, at a mean turn index of 2.0.
The same rooms under --cost-tiers, with the evidence-first opening.
arm correct % cost/ep correct/kU
vote 71.1 69.00 10.31
ladder 52.6 4.56 115.39
ladder+dir 84.4 3.24 260.33
hive+ 67.6 52.40 12.91
hive+cost 68.1 52.62 12.95
all-reasoning 67.6 115.27 5.87
hive+cost buys 12.95 right answers per thousand units against
all-reasoning's 5.87, and the whole of that gap is that all-reasoning puts
every seat on the ten-unit tier for the same 67.6%. Nothing about the
delegation mechanism produced it; not spending ten units on seats that do not
need them did.
5000 rooms, ±50 noise. This is the shape the mechanism was built for.
ordinary opening --blind-evidence
5000 rooms correct % fact % rho correct % 95% CI fact % knows % rho
ladder 35.1 — — 35.1 33.7–36.4 — — —
vote 15.0 — — 15.0 14.1–16.1 — — —
hive+ 15.3 1.5 0.74 66.3 65.0–67.6 96.8 0.0 0.37
hive+dir 15.3 1.5 0.74 65.8 64.5–67.1 95.7 77.5 0.42
hive+defer 15.4 1.4 0.74 66.8 65.4–68.1 96.8 0.0 0.07
hive+dir+defer 15.4 1.4 0.74 66.6 65.3–67.9 95.7 77.3 0.10
hive+ref 15.3 1.8 0.73 53.3 51.9–54.7 96.4 0.0 0.37
hive+ev 12.7 7.1 0.74 26.0 24.8–27.2 96.5 0.0 0.54
ladder+dir 34.6 — — 64.1 62.8–65.4 — 0.0 —
Under the ordinary opening no deliberating arm solves it, and fact % says
why: the deciding fact reaches the floor in time in 1.5% of episodes. A
!propose counts as a supporter, so four lay members who privately favour the
planted decoy carry it inside the blind round; the room is in Phase::Commit
before the fact-holder ever sees a floor to deposit against.
With the evidence-first opening fact % goes to 96.8% and the answer goes
from 15% to 66%. hive+ − vote is +51.3 [+49.8, +52.8], and it reproduces at
three seeds. That is a finding about when a member speaks, not about any
mechanism in the library: same rooms, same evaluations, same policy, same
fold, and a participant policy that states what it knows before what it wants.
On this shape hive+dir loses — 65.8% against hive+'s 66.3% — with
Knows winning the floor in 77.5% of episodes. The directory does not beat
the policy without it anywhere: +0.0 on the uniform room, +0.4 under the
evidence-first opening, +0.0 and +0.4 with two specialists, −0.5 here.
!defer is neutral by the same measure, moving ±0.5 and never leaving the
interval.
ladder+dir under --blind-evidence is an artifact. The arm tells its
router which topic the call turns on, and in this benchmark that topic is the
correct option; with an evidence-first opening the directory records who
deposited a reading of it, and a member who deposited on the truth usually
favours the truth. The 92-point swing from the same arm's 49.5% under the
ordinary opening is the size of the leak, not of any mechanism. The arm is kept,
unchanged and labelled, because deleting an arm that started scoring well would
be worse than explaining why its score is not evidence.
rho is the number the mechanism is judged by, and it separates the arms.
hive on a hidden profile under the ordinary opening reads 0.83 — the
directory reproducing the speaking order and having learned nothing except who
talked. Every deliberating arm of the uniform bench sits at 0.72–0.75.
Under the evidence-first opening it falls to 0.37, and with !defer on to
0.07, 0.06 at --defer-cap 2. Depositing before arguing earns
specialisation without earning a position; deferring zeroes a member's own
weight on a topic. Neither of those is the directory, and at 0.72 very little
of any result on the default bench could be credited to the estimate having
found something.
And history makes it worse. Over 2000 rooms with two specialists,
ladder+dir scores 52.1% at --history 0 — the null control, where every
description is None — then 44.8, 44.0, 43.9 and 43.6 at one, three, five and
ten prior episodes. It is not undertrained; it converges, slowly, on the wrong
member.
scenarios/index-lock-expert.txt is the hidden profile written for a room with
a named specialist, and scenarios/index-lock-tiers.txt is the same scenario
with a tier: on every seat so the routing can be priced. Polled alone against
a live model the scenario answers #rollback — three runs, two models,
plurality wrong every time and never more than one member on the truth, which
is the control being unable to win.
The matrix that is running:
| scenario | backend | repeats |
|---|---|---|
checkout-503 |
HTTP, flash
|
5 |
index-lock-expert |
HTTP, flash
|
5 |
index-lock-expert |
HTTP, --specialist-model reasoning
|
5 |
index-lock-expert |
claude -p --model flash |
3 |
index-lock-expert |
opencode run -m ladder/flash |
3 |
index-lock-expert |
codex exec against OpenRouter deepseek/deepseek-v4-flash
|
3 |
index-lock-tiers |
HTTP, mixed tiers | 5 |
index-lock-tiers |
HTTP, all-reasoning | 3 |
checkout-503-federated |
HTTP, flash
|
3 |
codex exec is a CLI row by necessity: the router cannot relay a streaming
Responses request, so it cannot go through the harness's own HTTP backend.
EpisodePolicy::DEFAULT carries directory: None and defer_cap: None. The
mechanism costs nothing on the uniform bench and buys nothing on the hidden
profile, which is a weaker case for defaulting it on rather than a stronger one.
--blind-evidence is off too, and for a sharper reason: it costs about seven
points on an ordinary room. It is a participant policy, not a library setting.
cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000
cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000 --blind-evidence
cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000 --specialists 2
cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000 \
--specialists 2 --blind-evidence --cost-tiers
cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000 --hidden-profile
cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000 \
--hidden-profile --blind-evidenceThe full write-up, every table and what the simulation does not show are in
docs/experiments/2026-09-05-expert-delegation.md,
and the harness's own notes are in
DELEGATION.md.
tinyhivemind is GPL-3.0-only. Built by @senamakel.
Start here
The algebra
- Shared medium
- Desks and rosters
- Mentions
- Cross-desk referral
- Transcript projection
- Threads
- Recall
- Private asides
- Responder ladder
Hive mechanics
- Hive episodes
- Trace grammar
- Transactive memory
- Episode policy
- Benchmarks
- Benchmarks: live rooms
- Benchmarks: delegation
- Benchmarks: private exchange
- Benchmarks: scale
Working on it
Reference