Skip to content

v3.2.1 — The bar held, the ruler grew

Latest

Choose a tag to compare

@sanlee-ys sanlee-ys released this 03 Aug 03:05
· 29 commits to main since this release
10f6348

The global-boundary prompt clause ships — the same clause v3.2.0-era ADR-023 measured and reverted at p=0.0522, re-run against a ruler twice the size and adopted at p=0.0002.

The bar never moved. The ruler did.

What changed

One bullet in the region rubric: naming a US service, command, program office, contractor, unit or official identifies the actor, not a place — so a story whose only geography is institutional is global, not americas. A second sentence bounds the fix, because ADR-020 is the measured precedent for what happens when restraint lives only in a prompt.

PATCH, not MINOR: the classifier is more correct; the {category, operational_domain, region} contract is untouched.

The measurement

All four pre-registered rules passed at effective n=595 (295 rows reused, 300 new, 5 exact duplicates dropped before grading), at a design power of 0.837.

Axis Baseline Candidate Lift cand/base McNemar p
region (target) 89.9% 94.1% +4.2 35 / 10 0.0002
category (guardrail) 93.4% 94.1% +0.7 11 / 7 0.4807
operational_domain (guardrail) 89.4% 92.9% +3.5 30 / 9 0.0011

20 of 32 named global pulls fixed, 8 correct rows dragged the other way, harness clean 595/595 on all three axes.

Published gold numbers (n=54, human-graded)

Adoption moved the prompt fingerprint a59689e8…b0202d06…, forcing a full paid re-run — workhorse and judge, because judge_region_agreement is a gated floor and had to be re-measured rather than inherited. This is the first release since v3.0.0 where these numbers move at all.

Gated metric Floor v3.2.0 v3.2.1
category_accuracy 0.83 92.6% 94.4%
category_macro_f1 0.85 0.911 0.930
domain_accuracy 0.83 92.6% 98.1%
domain_macro_f1 0.83 0.933 0.982
judge_category_agreement 0.83 92.6% 94.4%
judge_domain_agreement 0.88 98.1% 92.6%
region_accuracy 0.78 87.0% 94.4%
judge_region_agreement 0.93 100.0% 96.3%

All eight pass as written. None was moved, waived, or added. judge_region_agreement is the one with real authority: 2 judge-vs-human disagreements against a budget of 3, and the number that would have stopped this release.

On the gold set the region error mode inverted — 7 misses to 3, only one still the named cluster, the other two americas rows over-called to global. Five under-calls traded for two over-calls, and the collateral was priced before the run rather than discovered after it.

The honest lesson

ADR-023 ran at about 49% power against the effect it observed — a coin flip on whether it could detect its own effect. So its finding was underpowered, not refuted, and its verdict was correct on its date. It is amended with a dated pointer, not rewritten: honoring that threshold when it was inconvenient is the only reason this adoption means anything.

Measure-first is now seven-for-seven, and this is the first adoption after six declines.

Also in this release

  • The re-run harness goes dormant and refuses every run and report entry point, closing a quiet trap: after the adoption re-run the published gold record is produced by the clause prompt, so --gold-report would otherwise have compared the candidate arm against itself and printed a 0.0-point lift.
  • Published-marker cascade across architecture, portfolio, learning-notes and kb-agent.

Full detail: ADR-024 · CHANGELOG · reports in evals/region_clause_rerun.txt and evals/region_clause_gold_rerun.txt