Gate-2 amendment proposal 1: mean-over-draws scoring - #96
Conversation
…no self-rescue) Adds gate_2.amendment_proposed to gates.yaml as an inert public proposal object (mirrors gate 1's amendment-2 pattern). Per-cell scoring moves from one frozen simulation draw to the mean over K=20 draws per gate seed (draw seeds 5200+k), at the SAME locked tolerances and 4-of-5 conjunction. No locked value changes: gate_2.thresholds is byte-identical to origin/master (the diff is a pure insertion), locked stays true, and no model reads the block. Carries the full ceremony structure: changes with machine-bound derivations, operating-characteristics evidence recomputed from runs/gate2_forensics_v1 .json, honest disclosure (the estimator is not a pass-machine -- boundary level cells are not rescued; mean_lifetime_marriages|male rises 0.49->0.59), considered_and_rejected (widen tolerances / best-of-N / K=100), compute-cost note (~20 min/candidate), path_to_pass (fresh candidate 10), and an illustrative_retroactive_application block with applied:false marking the outer verdict not-computable (candidates 8/9 stay FAIL). Parallel gate-2 amendment test block in tests/test_gates_derivations.py binds the OC recompute, K=20 + seed convention, applied:false, the k=100 arithmetic and compute cost, and a byte compare of the locked thresholds vs origin/master. All dormant-safe (skip when amendment_proposed is absent). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
Verdict: AMEND BEFORE RATIFYING. Adversarial round on the gate-2 amendment-1 proposal object (head 1. Goalpost timing (check 1): accommodation-shaped on its face, but the estimator argument stands without candidates 8–9 — and the precedent's pressure-valve tell is absent. I verified the derivation-basis claim from 2. OC integrity (check 2): every table number recomputes. From 3. Selection hazards (check 3): closed, with one record-honesty gap — required fix B. The draw-seed convention is frozen and candidate-independent: the 20-draw design was registered at candidate-8 grading time (#42 comment 4913512779, item 3) and the 4. Contract coherence + candidate-10 spec (check 6): the flip's edit list, and three prospective gaps — required fix C. No locked language blocks the proposal as an inert object, but the flip PR must change, at minimum: 5. Scope, tests, and suite (checks 4+5): clean. Required fixes. The core change — mean over 20 pre-registered draws of the cell rate, at byte-identical tolerances, prospective-only — needs no redesign: it is the estimator the locked block's own ratified OC and the lock-round's "exactly the right null" claim already presupposed, it fails the triggering candidates harder, and every operating characteristic recomputes from committed artifacts. Fix the record; then ratify. |
A: c2 illustration outer_scores set to round(committed,4) from runs/gate2_hazard_v2.json (1.3593->1.359, 0.869->0.8689, 0.894->0.8938); c2 test tightened from abs=1e-3 to exact round(score,4), matching the c8/c9 convention. B: add currently_locked.registration_divergence (locked text says "simulation seed s"; all committed runs v1-v9 used default_rng(4200+seed) per protocol.sim_rng_rule); add flip_on_ratification.flip_edits enumerating the flip's locked-text edits (protocol.candidate, statistic, pass_rule r->rbar, faithful_candidate_oc basis note, stray DRAFT label); fix "byte -compares"->parsed-compare wording and relabel the 155.1s field as the 5-seed total. C: add fresh_run_artifact_schema (referenced from path_to_pass): per -draw per-cell rates (20x46x5) so rbar recomputes; a run -invalidating undefined-draw rule (no silent drop); a report-only per-draw dispersion disclosure (per-cell per-draw sd + max per-draw |ln| per cell). Amendment tests updated to bind A/B/C, dormant-safe (skip when amendment_proposed is absent). Locked gate_2.thresholds unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Referee fixes A/B/C appliedApplied the three required fixes from the adversarial review (comment 4915412987). The estimator proposal itself is unchanged; the locked A — the c2 illustration now recomputes to the last digit
B — record honesty
C — candidate-10 run contract (
|
|
Verdict: RATIFY AS-IS. Verification round on fix commit A — resolved: all three c2 digits recompute exactly, and the binding is now exact. I recomputed each
Note the fix corrected not just the flagged 1.3593 but also the other two rows (0.869 → 0.8689, 0.894 → 0.8938), which the tightened convention required — under the old B — resolved: the record now describes the locked rule as written, and the flip is enumerated.
C — resolved: the candidate-10 run contract is pinned. Cross-cutting — nothing new broke.
Round 1 found the estimator change itself sound and candidate-independent; this round confirms the record now matches it digit-for-digit. Fixes A/B/C verified resolved. RATIFY AS-IS — merge ratifies the proposal object; the live flip follows in its own authorized PR per |
Gate-2 amendment proposal 1: mean-over-draws scoring
A public amendment PROPOSAL for gate 2, mirroring gate 1's amendment-2 ceremony. It is an inert object:
gates.yamlgate_2 stayslocked: true, status: locked, and the entirethresholdssubtree is byte-identical toorigin/master(the change is a pure 671-line insertion of thegate_2.amendment_proposedsibling; a test byte-compares the subtree). No model reads it; the harness scores only the locked block. Nothing is ratified — this is the object a referee round adjudicates.The estimator change
Per-cell scoring moves from one frozen simulation draw to the mean over K=20 draws per gate seed:
|ln(r_candidate,s / rate_a,s)|,r_candidate,sfrom one replicate (sim_seed 4200+s).|ln(r̄_candidate,s / rate_a,s)|, wherer̄is the mean over K=20 draws of the cell rate (the mean of the candidate statistic across draws, not the mean of the |ln| scores). Draw seedsdefault_rng(5200+k), k=0..19 — the committed forensics convention (Gate-2 chronic-cell forensics: pathway, parity, and draw-noise decomposition #94).Why it's an estimator fix, not a loosening. The locked tolerances derive from the 100-seed half-vs-half split floor (
round(floor mean + 4·floor sd, 3)), which is real-vs-real and contains no simulation-draw noise. The single-draw estimator injects a per-cell draw-noise term the tolerance never budgeted for. Averaging over 20 pre-registered draws shrinks that injected noise by √20 and moves the certified statistic toward the draw-noise-free rate the tolerance was measured against — aligning the estimator with the tolerance's own derivation basis without touching the error budget.Operating characteristics (recomputed from the committed forensics #94)
Per failing cell: measured tilt vs tolerance, the single-draw clip probability the locked estimator carries, and the mean-of-20 estimator's. Single-draw column reproduces the artifact's committed
prob_train_draw_clips_toleranceto 1e-16; mean-of-20 is the same rate-scale normal model with sd/√20. Train-side (side B vs rate_b) — the only committed multi-draw evidence.The estimator is not a pass-machine. The four NOISE-DOMINATED cells collapse toward 0 (their single-draw clips were draw noise). The two BOUNDARY cells are not rescued:
share_widowed.75+|femalestays material, andmean_lifetime_marriages|malerises 0.49 → 0.59 because its systematic tilt genuinely exceeds tolerance on 4 of 5 seeds — averaging drives the estimator toward that real level, which fails harder. A model with gross level errors fails by orders of magnitude at any K: candidate 2's committed cells clip at 2.4–4.5× tolerance (e.g.share_widowed.65-74|femalescore 1.359 vs tol 0.300 on all 5 seeds), 15–136× the draw-noise sd.No self-rescue
Prospective-only (inherits
gate_2.governance.amendment_rules, which inherits gate 1). Candidates 1–9's committed verdicts stand (all FAIL). Theillustrative_retroactive_applicationblock isapplied: false: it examines candidates 8 and 9 (the triggering runs) and honestly marks the outer verdict NOT_COMPUTABLE_OUTER — the committed gate-2 artifacts hold only one outer draw per seed, and the forensics' 20 draws are train-side, so recomputing the amended outer verdict would need new outer simulations this proposal does not run. The train-side reading shows neither would pass anyway (each keeps a real boundary-level residual). Path to pass: a fresh candidate-10 registration, one-shot on seeds 0–4 under the amended estimator.Considered and rejected
Compute cost
One committed one-shot run (candidate 8) took 65.9 s for 5 seeds × 1 draw; K=20 multiplies the per-seed simulation ~20× → ≈ 22 min/candidate. The forensics measured 20 train-side draws × 5 seeds at 155.1 s per-seed compute — minutes, not hours. Recorded and acceptable.
Tests
A parallel gate-2 amendment block in
tests/test_gates_derivations.py(mirroring the gate-1 amendment tests) binds the proposal to the committed artifacts: per-cell OC recomputes fromruns/gate2_forensics_v1.json; K=20 and the 5200+k seed convention pinned; the k=100 arithmetic and compute cost recompute;applied: falsewith candidates 8/9 FAIL and the outer verdict marked not-computable; and a byte compare ofgate_2.thresholdsvsorigin/masterproving no locked value moved. All are dormant-safe (skip ifamendment_proposedis absent), exactly as the gate-1 proposal-object tests went dormant after ratification.Ceremony next steps
Adversarial referee round on this proposal → fixes if any → verification → maintainer ratification by merge → a follow-up that flips the estimator live in the locked protocol (as PR #79/#81 did for the gate-2 lock and #67/#69 for gate 1's amendment 2). Do not merge this PR as a ratification.
Evidence chain: #94 (forensics) + #93 (candidate 8) / #95 (candidate 9) runs + the candidate 1–9 ladder on #42.
🤖 Generated with Claude Code