Flip the ratified gate-1 amendment live - #59
Merged
Conversation
The 2026-07-06 amendment (proposed, adversarially refereed, and ratified by the merge of PR 57, commit 0c324a7) now lives in the locked block: the runs view gates coverage only, with c2st_auc reported-not-gated and annotated; the benefit_space block sits under thresholds with its metrics, derivations, provenance, disclosures, and the amended seed-conjunction pass rule; amendment_proposed is replaced by an amendment_history record. Folds in the referee's two deferred nits (the builder docstring path; the demotion removing the runs c2st from geometry and derivations together, preserving the set-equality test). Nine historical run artifacts pin the thresholds they ran against; their consistency tests are now amendment-aware — a stored metric that a ratified amendment later demoted (per the view's reported_not_gated list) is excluded from the equality, so the artifacts remain the correct record of the gate as run. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
MaxGhenis
added a commit
that referenced
this pull request
Jul 7, 2026
Reported, not gated; no holdout-real contact beyond the pooled Q0 statistic the ratified benefit-space gate (PR #59) already scores. Reads no gate, changes no gate; writes only runs/q0_forensics_v1.json. Localizes the candidate-7 Q0 PIA-proxy overstatement (+9.30% pooled, PR #56) mechanically, so candidate 9's Q0 component is designed against the cause rather than against candidate 8's failed age+observed-span conditioning (PR #58, +12.2%). Regenerates candidate 7 (and, as a cross-check, candidate 8) deterministically over each gate seed via the merged machinery and compares against TRAIN zero-anchor persons' real careers (seed 0 pinned to float precision). Ranked mechanisms (pooled across seeds 0-4), exact additive decomposition of the weighted-mean Q0 PIA-proxy gap: participation (net zero<->positive): +10.12pp <- the whole gap zero->positive (never-workers resurrected): +20.68 positive->zero: -10.55 positive-earnings level (top-N average): -0.83pp <- not the cause An independent rank-transport counterfactual agrees: candidate participation pattern on the real Q0 level distribution reproduces +10.02% (the full gap); real participation on candidate levels gives +1.17% (~zero). The PIA-proxy averages the top min(10,n) positive years, so extra positive years inflate it only by crossing a never-worker from PIA=0 to positive; the top-N selection is inert (2% of Q0 persons have n_pos>10). Why: the RegimeGatedQRF sign gate, fit on all train pairs, overshoots the zero->positive re-entry rate weakly-attached persons actually have. Generated Q0 share-all-zero 0.240 vs train-real 0.297 (loses 5.7pp of never-workers); generated mean n_pos 3.47 vs 2.95. The re-entry rank pool compounds it slightly: it draws mean rank 0.302 vs the honest zero-anchor pre-gap 0.275 (+0.027) because ~39% of its weight is attached persons at higher pre-gap ranks, conditioned only on the near-constant u_A=p0/2. What would discriminate: among train-real zero-anchor persons, production-available covariates explain little of career-PIA variance (age+span R^2 0.131; adding anchor position -> 0.135), while realized n_pos alone gives 0.316. Candidate 8 used exactly age+span for Q0 and, reusing candidate 7's participation gate byte-for-byte (generated all-zero share identical every seed), moved only the LEVEL channel and the wrong way (level -0.83 -> +2.95), so its total gap grew to +12.16%. Design implication for candidate 9: fix the participation LAW for zero-anchor persons, not its conditioning. The gate must down-weight zero->positive re-entry to match the train-real zero-anchor attachment rate (share-all-zero and mean n_pos), and the re-entry rank pool should condition on zero-anchor donors so the near-constant u_A stops importing attached persons' higher pre-gap ranks. Leave the level machinery alone; it is already inside the noise floor. runs/q0_forensics_v1.json carries all per-seed and pooled numbers plus the candidate-8 cross-check; tests/test_q0_forensics.py pins the decomposition arithmetic and reproduces seed 0 live. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
MaxGhenis
added a commit
that referenced
this pull request
Jul 7, 2026
… regime (fails under the amended gate) (#62) The eleventh pre-registered gate-1 run and the first candidate scored under the amended gate (PR #57/#59). Candidate 7's machinery verbatim with two registered changes; the only calibration is the registered train-side SMM for lambda. Verdict under the amended gate: FAIL. Geometry 0/5, battery 2/5, pooled Q0 -17.89% (> 5). Two registered changes: 1. SMM-calibrated donor-coordinate blend. The k-NN third distance term becomes |lambda*u_w(donor) + (1-lambda)*u_A(donor) - u_A(target)| at the 0.25 weight; u_w is candidate 8's shrunk permanent rank, u_A the anchor ranks. lambda is calibrated per seed on the train split by SMM in the 5b tradition (grid {0,...,1.0}, autocorrelation-ladder SSE on the first 2,000 train persons, ties to the smaller lambda). lambda=0 reproduces candidate 7; lambda=1 candidate 8. 2. Zero-anchor participation regime. Zero-anchor holdout persons draw participation from a gate refit only on zero-anchor train pairs and their re-entry innovations from a zero-anchor-restricted re-entry pool. Positive-anchor persons keep the shared gate and full pools exactly as candidate 7. Candidate 9 does NOT adopt candidate 8's attachment distance. Findings: - Chosen lambda per seed: {0: 0.2, 1: 0.0, 2: 0.3, 3: 0.1, 4: 0.0}. The train SMM never chose lambda above 0.3 and chose lambda=0 (candidate 7) on two seeds. - The c9 pooled 10-year autocorrelation rung (0.514) lands inside the reference band [0.469, 0.609] and between the c7/c8 bracket (0.459/0.670) -- the blend achieved its aim on the 10-year rung in the pooled mean. But no single lambda on the grid lands all three rungs simultaneously: the registered risk materialized (the blend changes all three rungs together). Battery passes only 2/5 (seeds 2, 3); seed 0 fails the 2-year rung (dev 0.059), seeds 1/4 (lambda=0) leave the 10-year rung short. - Geometry 0/5: the binding constraint is the pairs-view c2st_auc (0.531-0.550, all just over 0.53) on all five seeds; benefit-space additionally fails on seeds 0/2/3. - The zero-anchor participation regime closed the never-worker resurrection (generated all-zero share 0.3155 vs real 0.3147, gap +0.0007, vs PR #61's +20.7pp for the shared gate) -- the participation law now matches reality. But the level over-corrected: pooled Q0 PIA-proxy moved from c7 +9.3% / c8 +12.2% to c9 -17.89% (the restricted re-entry pool plus the resurrection fix together subtract more than the +9.3% they were meant to remove). Artifact runs/gate1_rank_knn_v3.json (schema gate1_rank_knn.v3): per-seed chosen lambda + SMM ladders, the amended-gate scorecard (the benefit_space block per seed + pooled Q0), Q0 participation diagnostics, and the standard diagnostics. The battery-reference bit-exact precheck reproduced every committed value before scoring. Tests: seed-0 reproduction (live in .venv-gate) + the amended-verdict recomputation block (24 pass in .venv-gate; 311 pass / 16 skip in the repo .venv). Links issue #42; base machinery candidate 7 (#55); u_w candidate 8 (#58); benefit-space functional #56; C2ST forensics #54; amended gate #57/#59; Q0 forensics #61. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
MaxGhenis
added a commit
that referenced
this pull request
Jul 7, 2026
…try under the amended gate) (#64) The twelfth pre-registered gate-1 run, and the first whose every constant was selected by nested validation (PR #63) rather than by outer-gate feedback. Candidate 10 is the inner sweep's V1-lam0.1 variant at outer scale: candidate 7's k-NN conditional-rank-bootstrap machinery, candidate 9's zero-anchor participation regime with its two poisons removed (the re-entry-pool restriction dropped, Q0 targets held memory-exempt), and a FIXED lambda = 0.1 donor-coordinate blend for the non-Q0 targets. No calibration stage: lambda is fixed, not chosen against any score. Frozen spec: issue #42 comment 4902561460. Scored under the amended gate (gates.yaml gate_1, PR #57/#59) exactly as run 11: runs-view c2st demoted, gated benefit-space block folded into the geometry verdict, >=4/5 geometry AND >=4/5 battery AND pooled Q0 (abs <= 5). Verdict, exactly as computed: gate_1_pass = FALSE -- geometry 3/5 (needs 4/5), battery 4/5 (passes), pooled Q0 +0.038% (passes). Seeds 0 and 1 each clip the pairs-view C2ST (0.5315, 0.5330 vs 0.53); seeds 2/3/4 clear it. The only battery failure is seed 0's 2-year autocorrelation (0.7896, deviation 0.0595 > 0.05) -- a short-lag overshoot from the memory injection, not the 10-year undershoot of earlier candidates (every seed clears 4yr and 10yr). Inner-vs-outer: the ~0.005-0.01 heat correction the inner sweep predicted for the pairs C2ST delivered a reshuffle, not a uniform cooling -- mean cooled -0.0014 (inner 0.5293 -> outer 0.5279), but seeds 0/1 heated up (crossing to fail) while the worst inner seed (3, 0.5365) cooled to a comfortable pass (0.5260). Still 3/5, one seed short -- exactly the forecast's named failure (comment 4902561584, P(pass) ~0.42). The Q0 program was solved: pooled Q0 generalized inner +1.19% -> outer +0.04%, and the generated-vs-real all-zero share gap is +0.0007 (no never-worker resurrection). Generation verified byte-identical to the inner sweep's generate_variant at memory_mode=lambda_blend, lam=0.1, use_zero_anchor_gate=True under candidate-7 seeding. Seed-0 reproduction run live in .venv-gate (float-exact). Full pytest green in the repo .venv (361 passed, 22 skipped). ruff + black -l 79 clean. Deliverables: - scripts/run_gate1_candidate10.py -- deterministic; candidate-7 two-element substreams; the registered rules exactly. - runs/gate1_rank_knn_v4.json -- schema gate1_rank_knn.v4; spec_registration = the candidate-10 comment; standard diagnostics + Q0 participation numbers + per-seed benefit_space + the amended scorecard. - tests/test_gate1_qrf_candidate10.py -- seed-0 reproduction (skipif PSID + importorskip populace.fit) + generation-equivalence to the inner sweep + the Q0-exempt/full-pool checks + the amended-gate consistency block. Links: issue #42; PRs #63 (inner-validation harness + design sweep), #62 (candidate 9), #61 (Q0 forensics), #55 (candidate 7). Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This was referenced Jul 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The pre-agreed follow-up to #57 (ratified by merge, 0c324a7), mirroring the #33 → #39 pattern: the amendment content moves from the proposal subsection into the locked block. Runs-view c2st_auc → reported-not-gated (annotated on the view; disclosure retained); the gated benefit_space block and the amended pass rule go live; amendment_history records the full ceremony (review 4898105589 → fixes d02cdf5 → verification 4898186646 → ratification 0c324a7). Includes the referee's two deferred nits and makes the nine historical artifacts' consistency tests amendment-aware (stored thresholds remain the correct record of the gate as run). Full suite 290 passed / 13 skipped.
🤖 Generated with Claude Code