The global-boundary prompt clause ships — the same clause v3.2.0-era ADR-023 measured and reverted at p=0.0522, re-run against a ruler twice the size and adopted at p=0.0002.
The bar never moved. The ruler did.
What changed
One bullet in the region rubric: naming a US service, command, program office, contractor, unit or official identifies the actor, not a place — so a story whose only geography is institutional is global, not americas. A second sentence bounds the fix, because ADR-020 is the measured precedent for what happens when restraint lives only in a prompt.
PATCH, not MINOR: the classifier is more correct; the {category, operational_domain, region} contract is untouched.
The measurement
All four pre-registered rules passed at effective n=595 (295 rows reused, 300 new, 5 exact duplicates dropped before grading), at a design power of 0.837.
| Axis | Baseline | Candidate | Lift | cand/base | McNemar p |
|---|---|---|---|---|---|
| region (target) | 89.9% | 94.1% | +4.2 | 35 / 10 | 0.0002 |
| category (guardrail) | 93.4% | 94.1% | +0.7 | 11 / 7 | 0.4807 |
| operational_domain (guardrail) | 89.4% | 92.9% | +3.5 | 30 / 9 | 0.0011 |
20 of 32 named global pulls fixed, 8 correct rows dragged the other way, harness clean 595/595 on all three axes.
Published gold numbers (n=54, human-graded)
Adoption moved the prompt fingerprint a59689e8… → b0202d06…, forcing a full paid re-run — workhorse and judge, because judge_region_agreement is a gated floor and had to be re-measured rather than inherited. This is the first release since v3.0.0 where these numbers move at all.
| Gated metric | Floor | v3.2.0 | v3.2.1 |
|---|---|---|---|
category_accuracy |
0.83 | 92.6% | 94.4% |
category_macro_f1 |
0.85 | 0.911 | 0.930 |
domain_accuracy |
0.83 | 92.6% | 98.1% |
domain_macro_f1 |
0.83 | 0.933 | 0.982 |
judge_category_agreement |
0.83 | 92.6% | 94.4% |
judge_domain_agreement |
0.88 | 98.1% | 92.6% |
region_accuracy |
0.78 | 87.0% | 94.4% |
judge_region_agreement |
0.93 | 100.0% | 96.3% |
All eight pass as written. None was moved, waived, or added. judge_region_agreement is the one with real authority: 2 judge-vs-human disagreements against a budget of 3, and the number that would have stopped this release.
On the gold set the region error mode inverted — 7 misses to 3, only one still the named cluster, the other two americas rows over-called to global. Five under-calls traded for two over-calls, and the collateral was priced before the run rather than discovered after it.
The honest lesson
ADR-023 ran at about 49% power against the effect it observed — a coin flip on whether it could detect its own effect. So its finding was underpowered, not refuted, and its verdict was correct on its date. It is amended with a dated pointer, not rewritten: honoring that threshold when it was inconvenient is the only reason this adoption means anything.
Measure-first is now seven-for-seven, and this is the first adoption after six declines.
Also in this release
- The re-run harness goes dormant and refuses every run and report entry point, closing a quiet trap: after the adoption re-run the published gold record is produced by the clause prompt, so
--gold-reportwould otherwise have compared the candidate arm against itself and printed a 0.0-point lift. - Published-marker cascade across architecture, portfolio, learning-notes and kb-agent.
Full detail: ADR-024 · CHANGELOG · reports in evals/region_clause_rerun.txt and evals/region_clause_gold_rerun.txt