Skip to content

Benchmarks delegation

Steven Enamakel edited this page Sep 5, 2026 · 6 revisions

Benchmarks: delegation and expertise

The delegation half of the deliberation benchmark: what happens when a room's members are not interchangeable — one of them holds the fact that decides the question, and the others do not.

It answers three questions, written down before any of these numbers existed.

  • Q1 — does the floor reach the member who holds the fact? fact % is the share of episodes in which the decisive member deposited a topiced !evidence line before the commit boundary, and to-fact is the mean turn index at which it did. A deposit landing at or after the boundary is compute the room paid for and could not use, and is scored as a miss.
  • Q2 — does an informative router beat a uniform ladder? route %, over the two ladder arms.
  • Q3 — what does accuracy cost, once seats have prices? cost/ep and correct/kU, right answers per thousand units, under --cost-tiers.

A fourth number, rho, is an obligation rather than a question: the spec requires this benchmark to publish the rank correlation between a member's directory weight and its share of the episode's turns, and to report the mechanism as having failed if it merely tracks who talked.

The flags

flag what it does
--specialists N N members read one topic each far more tightly than everybody else, and everybody else's read of that topic widens to match. Information is redistributed, not created.
--hidden-profile One decoy is planted above every member's own argmax except one member's, and that member alone holds the fact that rules it out.
--blind-evidence A member's first turn, while the room is blind, is a deposit rather than a position. Off by default.
--defer-cap N Turns a member may spend saying a topic is not theirs.
--cost-tiers A specialist's turn costs ten units against a lay member's one.
--history N Prior episodes of hive+ from which ladder+dir earns its directory.

The arms

hive+dir is the tuned policy with the directory on, so BidReason::Knows is reachable. hive+defer adds defer_cap with nothing routing the vacated turn, and hive+dir+defer is both. ladder+dir is the responder ladder with each candidate's earned directory lines rendered into its description and the contested topic named in the request — the routing rule a subagent description and a role string implement — validated through the real accept_selection. hive+cost and all-reasoning appear under --cost-tiers only.

The uniform room, which predicted zero

5000 rooms, ±90 noise. With homogeneous expertise there is nothing to route on, and the spec predicted exactly that.

                      ordinary opening        --blind-evidence
5000 rooms          correct %      95% CI   correct %      95% CI  knows %
ladder                   57.6   56.2–59.0        57.6   56.2–59.0        —
vote                     78.5   77.4–79.6        78.5   77.4–79.6        —
hive+                    82.1   81.0–83.1        75.3   74.1–76.5      0.0
hive+dir                 82.1   81.0–83.1        75.7   74.5–76.8     75.1
hive+defer               82.1   81.0–83.1        75.3   74.1–76.5      0.0
hive+dir+defer           82.1   81.0–83.1        75.7   74.5–76.8     75.1
ladder+dir               49.5   48.1–50.9        98.8   98.5–99.1      0.0

Under the ordinary opening the delegation arms are hive+ to the digit and knows % is 0.0: BidReason::Knows never fires. It needs a member that is the directory's top holder of the contested topic and has taken no position on it, and where every member opens with a !propose those never hold together. ladder+dir is eight points worse than the uninformed draw over the same five candidates.

Two specialists

                      ordinary opening        --blind-evidence
5000 rooms          correct %      95% CI   correct %      95% CI  knows %  route %
ladder                   52.6   51.2–54.0        52.6   51.2–54.0        —     18.9
vote                     71.1   69.9–72.4        71.1   69.9–72.4        —        —
hive+                    74.2   72.9–75.4        67.6   66.3–68.9      0.0        —
hive+dir                 74.2   72.9–75.4        68.0   66.7–69.2     78.3        —
hive+defer               74.3   73.1–75.5        67.2   65.9–68.5      0.0        —
hive+dir+defer           74.3   73.1–75.5        68.1   66.8–69.4     79.8        —
ladder+dir               45.1   43.8–46.5        84.4   83.4–85.4      0.0     22.3

ladder+dir routes to the decisive member more often than ladder does — 22.3% against 18.9% — and is seven and a half points less accurate. Directory weight on a topic is earned by grounding it, so the heaviest holder is the member who argued it hardest rather than the one who reads it best. Routing precisely to the wrong criterion is worse than not routing.

Under --blind-evidence the specialist deposits before the commit boundary in 93.4% of episodes, at a mean turn index of 2.0.

The price of a seat

The same rooms under --cost-tiers, with the evidence-first opening.

arm              correct %   cost/ep    correct/kU
vote                  71.1     69.00         10.31
ladder                52.6      4.56        115.39
ladder+dir            84.4      3.24        260.33
hive+                 67.6     52.40         12.91
hive+cost             68.1     52.62         12.95
all-reasoning         67.6    115.27          5.87

hive+cost buys 12.95 right answers per thousand units against all-reasoning's 5.87, and the whole of that gap is that all-reasoning puts every seat on the ten-unit tier for the same 67.6%. Nothing about the delegation mechanism produced it; not spending ten units on seats that do not need them did.

The hidden profile

5000 rooms, ±50 noise. This is the shape the mechanism was built for.

                      ordinary opening                --blind-evidence
5000 rooms          correct %   fact %    rho   correct %      95% CI   fact %  knows %    rho
ladder                   35.1        —      —        35.1   33.7–36.4        —        —      —
vote                     15.0        —      —        15.0   14.1–16.1        —        —      —
hive+                    15.3      1.5   0.74        66.3   65.0–67.6     96.8      0.0   0.37
hive+dir                 15.3      1.5   0.74        65.8   64.5–67.1     95.7     77.5   0.42
hive+defer               15.4      1.4   0.74        66.8   65.4–68.1     96.8      0.0   0.07
hive+dir+defer           15.4      1.4   0.74        66.6   65.3–67.9     95.7     77.3   0.10
hive+ref                 15.3      1.8   0.73        53.3   51.9–54.7     96.4      0.0   0.37
hive+ev                  12.7      7.1   0.74        26.0   24.8–27.2     96.5      0.0   0.54
ladder+dir               34.6        —      —        64.1   62.8–65.4        —      0.0      —

Under the ordinary opening no deliberating arm solves it, and fact % says why: the deciding fact reaches the floor in time in 1.5% of episodes. A !propose counts as a supporter, so four lay members who privately favour the planted decoy carry it inside the blind round; the room is in Phase::Commit before the fact-holder ever sees a floor to deposit against.

With the evidence-first opening fact % goes to 96.8% and the answer goes from 15% to 66%. hive+ − vote is +51.3 [+49.8, +52.8], and it reproduces at three seeds. That is a finding about when a member speaks, not about any mechanism in the library: same rooms, same evaluations, same policy, same fold, and a participant policy that states what it knows before what it wants.

On this shape hive+dir loses — 65.8% against hive+'s 66.3% — with Knows winning the floor in 77.5% of episodes. The directory does not beat the policy without it anywhere: +0.0 on the uniform room, +0.4 under the evidence-first opening, +0.0 and +0.4 with two specialists, −0.5 here. !defer is neutral by the same measure, moving ±0.5 and never leaving the interval.

Two circularity notes

ladder+dir under --blind-evidence is an artifact. The arm tells its router which topic the call turns on, and in this benchmark that topic is the correct option; with an evidence-first opening the directory records who deposited a reading of it, and a member who deposited on the truth usually favours the truth. The 92-point swing from the same arm's 49.5% under the ordinary opening is the size of the leak, not of any mechanism. The arm is kept, unchanged and labelled, because deleting an arm that started scoring well would be worse than explaining why its score is not evidence.

rho is the number the mechanism is judged by, and it separates the arms. hive on a hidden profile under the ordinary opening reads 0.83 — the directory reproducing the speaking order and having learned nothing except who talked. Every deliberating arm of the uniform bench sits at 0.720.75. Under the evidence-first opening it falls to 0.37, and with !defer on to 0.07, 0.06 at --defer-cap 2. Depositing before arguing earns specialisation without earning a position; deferring zeroes a member's own weight on a topic. Neither of those is the directory, and at 0.72 very little of any result on the default bench could be credited to the estimate having found something.

And history makes it worse. Over 2000 rooms with two specialists, ladder+dir scores 52.1% at --history 0 — the null control, where every description is None — then 44.8, 44.0, 43.9 and 43.6 at one, three, five and ten prior episodes. It is not undertrained; it converges, slowly, on the wrong member.

Live rooms

scenarios/index-lock-expert.txt is the hidden profile written for a room with a named specialist, and scenarios/index-lock-tiers.txt is the same scenario with a tier: on every seat so the routing can be priced. Polled alone against a live model the scenario answers #rollback — three runs, two models, plurality wrong every time and never more than one member on the truth, which is the control being unable to win.

The matrix that is running:

scenario backend repeats
checkout-503 HTTP, flash 5
index-lock-expert HTTP, flash 5
index-lock-expert HTTP, --specialist-model reasoning 5
index-lock-expert claude -p --model flash 3
index-lock-expert opencode run -m ladder/flash 3
index-lock-expert codex exec against OpenRouter deepseek/deepseek-v4-flash 3
index-lock-tiers HTTP, mixed tiers 5
index-lock-tiers HTTP, all-reasoning 3
checkout-503-federated HTTP, flash 3

codex exec is a CLI row by necessity: the router cannot relay a streaming Responses request, so it cannot go through the harness's own HTTP backend.

Why both knobs ship off

EpisodePolicy::DEFAULT carries directory: None and defer_cap: None. The mechanism costs nothing on the uniform bench and buys nothing on the hidden profile, which is a weaker case for defaulting it on rather than a stronger one. --blind-evidence is off too, and for a sharper reason: it costs about seven points on an ordinary room. It is a participant policy, not a library setting.

Reproducing

cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000
cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000 --blind-evidence
cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000 --specialists 2
cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000 \
  --specialists 2 --blind-evidence --cost-tiers
cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000 --hidden-profile
cargo run --release -p tinyhivemind-hive --example bench -- --episodes 5000 \
  --hidden-profile --blind-evidence

The full write-up, every table and what the simulation does not show are in docs/experiments/2026-09-05-expert-delegation.md, and the harness's own notes are in DELEGATION.md.

Clone this wiki locally