There was an error while loading. Please reload this page.
Do not fabricate corrected early-expert counts; add an honest caveat Codex: capping the pre-fix commit-boundary approximation at each row's decided count assumed the two definitions agree on every decided round, but a decided round whose commit phase spanned multiple turns can flip under the corrected metric too -- the assumption was unverified. Reverted to the originally recorded figures and added a caveat explaining what changed and why it cannot be resolved without the raw per-round turn indices, which this report does not retain. CodeRabbit: clarified that expert reports two independent counts (before / cited), not a ratio. Co-authored-by: Medulla <medulla@tinyhumans.ai>
Recompute the corrected early-expert totals and cost-ratio wording Codex: the print_expert fix in the paired PR proves before cannot exceed a row's decided count; capped the three affected rows (flash, opencode, codex exec) from 3 to their decided count of 2, which brings the derived total to a clean 20 of 20 decided rounds rather than 23 of 23 rounds. Co-authored-by: Medulla <medulla@tinyhumans.ai>
Record the corrected mixed-tier live row and fix the harness-defect narrative The tier-only specialist seating fix in the paired tinyhivemind PR replaces the buggy --specialist-model row (which seated the whole room on reasoning) with a genuine mixed-tier measurement: only dba on reasoning, the other four seats on flash. Recompute the matrix totals with the new row in place. Co-authored-by: Medulla <medulla@tinyhumans.ai>
Record the live delegation matrix on the benchmark pages Co-authored-by: Medulla <medulla@tinyhumans.ai>
Refresh rho values for the exact Pearson-on-ranks statistic The bench's Spearman shortcut returned a spurious 1.0 for fully tied vectors; every rho value here is regenerated against the fixed exact integer Pearson-on-tie-averaged-ranks statistic. Non-rho numbers were re-checked against fresh runs and are unchanged. Co-authored-by: Medulla <medulla@tinyhumans.ai>
Record the delegation benchmark and trim Benchmarks below the cap Co-authored-by: Medulla <medulla@tinyhumans.ai>