Skip to content

Flip the ratified gate-1 amendment live - #59

Merged
MaxGhenis merged 1 commit into
masterfrom
gate1-amendment-flip
Jul 6, 2026
Merged

Flip the ratified gate-1 amendment live#59
MaxGhenis merged 1 commit into
masterfrom
gate1-amendment-flip

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

The pre-agreed follow-up to #57 (ratified by merge, 0c324a7), mirroring the #33#39 pattern: the amendment content moves from the proposal subsection into the locked block. Runs-view c2st_auc → reported-not-gated (annotated on the view; disclosure retained); the gated benefit_space block and the amended pass rule go live; amendment_history records the full ceremony (review 4898105589 → fixes d02cdf5 → verification 4898186646 → ratification 0c324a7). Includes the referee's two deferred nits and makes the nine historical artifacts' consistency tests amendment-aware (stored thresholds remain the correct record of the gate as run). Full suite 290 passed / 13 skipped.

🤖 Generated with Claude Code

The 2026-07-06 amendment (proposed, adversarially refereed, and
ratified by the merge of PR 57, commit 0c324a7) now lives in the
locked block: the runs view gates coverage only, with c2st_auc
reported-not-gated and annotated; the benefit_space block sits under
thresholds with its metrics, derivations, provenance, disclosures,
and the amended seed-conjunction pass rule; amendment_proposed is
replaced by an amendment_history record. Folds in the referee's two
deferred nits (the builder docstring path; the demotion removing the
runs c2st from geometry and derivations together, preserving the
set-equality test).

Nine historical run artifacts pin the thresholds they ran against;
their consistency tests are now amendment-aware — a stored metric
that a ratified amendment later demoted (per the view's
reported_not_gated list) is excluded from the equality, so the
artifacts remain the correct record of the gate as run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@vercel

vercel Bot commented Jul 6, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
social-security-model Ready Ready Preview, Comment Jul 6, 2026 11:37pm

Request Review

@MaxGhenis
MaxGhenis merged commit 940f308 into master Jul 6, 2026
7 checks passed
MaxGhenis added a commit that referenced this pull request Jul 7, 2026
Reported, not gated; no holdout-real contact beyond the pooled Q0
statistic the ratified benefit-space gate (PR #59) already scores. Reads
no gate, changes no gate; writes only runs/q0_forensics_v1.json.

Localizes the candidate-7 Q0 PIA-proxy overstatement (+9.30% pooled, PR
#56) mechanically, so candidate 9's Q0 component is designed against the
cause rather than against candidate 8's failed age+observed-span
conditioning (PR #58, +12.2%). Regenerates candidate 7 (and, as a
cross-check, candidate 8) deterministically over each gate seed via the
merged machinery and compares against TRAIN zero-anchor persons' real
careers (seed 0 pinned to float precision).

Ranked mechanisms (pooled across seeds 0-4), exact additive
decomposition of the weighted-mean Q0 PIA-proxy gap:

  participation (net zero<->positive): +10.12pp  <- the whole gap
    zero->positive (never-workers resurrected): +20.68
    positive->zero:                             -10.55
  positive-earnings level (top-N average):  -0.83pp  <- not the cause

An independent rank-transport counterfactual agrees: candidate
participation pattern on the real Q0 level distribution reproduces
+10.02% (the full gap); real participation on candidate levels gives
+1.17% (~zero). The PIA-proxy averages the top min(10,n) positive years,
so extra positive years inflate it only by crossing a never-worker from
PIA=0 to positive; the top-N selection is inert (2% of Q0 persons have
n_pos>10).

Why: the RegimeGatedQRF sign gate, fit on all train pairs, overshoots
the zero->positive re-entry rate weakly-attached persons actually have.
Generated Q0 share-all-zero 0.240 vs train-real 0.297 (loses 5.7pp of
never-workers); generated mean n_pos 3.47 vs 2.95. The re-entry rank
pool compounds it slightly: it draws mean rank 0.302 vs the honest
zero-anchor pre-gap 0.275 (+0.027) because ~39% of its weight is
attached persons at higher pre-gap ranks, conditioned only on the
near-constant u_A=p0/2.

What would discriminate: among train-real zero-anchor persons,
production-available covariates explain little of career-PIA variance
(age+span R^2 0.131; adding anchor position -> 0.135), while realized
n_pos alone gives 0.316. Candidate 8 used exactly age+span for Q0 and,
reusing candidate 7's participation gate byte-for-byte (generated
all-zero share identical every seed), moved only the LEVEL channel and
the wrong way (level -0.83 -> +2.95), so its total gap grew to +12.16%.

Design implication for candidate 9: fix the participation LAW for
zero-anchor persons, not its conditioning. The gate must down-weight
zero->positive re-entry to match the train-real zero-anchor attachment
rate (share-all-zero and mean n_pos), and the re-entry rank pool should
condition on zero-anchor donors so the near-constant u_A stops importing
attached persons' higher pre-gap ranks. Leave the level machinery alone;
it is already inside the noise floor.

runs/q0_forensics_v1.json carries all per-seed and pooled numbers plus
the candidate-8 cross-check; tests/test_q0_forensics.py pins the
decomposition arithmetic and reproduces seed 0 live.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Jul 7, 2026
… regime (fails under the amended gate) (#62)

The eleventh pre-registered gate-1 run and the first candidate scored
under the amended gate (PR #57/#59). Candidate 7's machinery verbatim
with two registered changes; the only calibration is the registered
train-side SMM for lambda.

Verdict under the amended gate: FAIL. Geometry 0/5, battery 2/5, pooled
Q0 -17.89% (> 5).

Two registered changes:
1. SMM-calibrated donor-coordinate blend. The k-NN third distance term
   becomes |lambda*u_w(donor) + (1-lambda)*u_A(donor) - u_A(target)| at
   the 0.25 weight; u_w is candidate 8's shrunk permanent rank, u_A the
   anchor ranks. lambda is calibrated per seed on the train split by SMM
   in the 5b tradition (grid {0,...,1.0}, autocorrelation-ladder SSE on
   the first 2,000 train persons, ties to the smaller lambda). lambda=0
   reproduces candidate 7; lambda=1 candidate 8.
2. Zero-anchor participation regime. Zero-anchor holdout persons draw
   participation from a gate refit only on zero-anchor train pairs and
   their re-entry innovations from a zero-anchor-restricted re-entry
   pool. Positive-anchor persons keep the shared gate and full pools
   exactly as candidate 7. Candidate 9 does NOT adopt candidate 8's
   attachment distance.

Findings:
- Chosen lambda per seed: {0: 0.2, 1: 0.0, 2: 0.3, 3: 0.1, 4: 0.0}. The
  train SMM never chose lambda above 0.3 and chose lambda=0 (candidate 7)
  on two seeds.
- The c9 pooled 10-year autocorrelation rung (0.514) lands inside the
  reference band [0.469, 0.609] and between the c7/c8 bracket
  (0.459/0.670) -- the blend achieved its aim on the 10-year rung in the
  pooled mean. But no single lambda on the grid lands all three rungs
  simultaneously: the registered risk materialized (the blend changes
  all three rungs together). Battery passes only 2/5 (seeds 2, 3);
  seed 0 fails the 2-year rung (dev 0.059), seeds 1/4 (lambda=0) leave
  the 10-year rung short.
- Geometry 0/5: the binding constraint is the pairs-view c2st_auc
  (0.531-0.550, all just over 0.53) on all five seeds; benefit-space
  additionally fails on seeds 0/2/3.
- The zero-anchor participation regime closed the never-worker
  resurrection (generated all-zero share 0.3155 vs real 0.3147, gap
  +0.0007, vs PR #61's +20.7pp for the shared gate) -- the participation
  law now matches reality. But the level over-corrected: pooled Q0
  PIA-proxy moved from c7 +9.3% / c8 +12.2% to c9 -17.89% (the restricted
  re-entry pool plus the resurrection fix together subtract more than the
  +9.3% they were meant to remove).

Artifact runs/gate1_rank_knn_v3.json (schema gate1_rank_knn.v3):
per-seed chosen lambda + SMM ladders, the amended-gate scorecard (the
benefit_space block per seed + pooled Q0), Q0 participation diagnostics,
and the standard diagnostics. The battery-reference bit-exact precheck
reproduced every committed value before scoring. Tests: seed-0
reproduction (live in .venv-gate) + the amended-verdict recomputation
block (24 pass in .venv-gate; 311 pass / 16 skip in the repo .venv).

Links issue #42; base machinery candidate 7 (#55); u_w candidate 8
(#58); benefit-space functional #56; C2ST forensics #54; amended gate
#57/#59; Q0 forensics #61.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Jul 7, 2026
…try under the amended gate) (#64)

The twelfth pre-registered gate-1 run, and the first whose every constant was
selected by nested validation (PR #63) rather than by outer-gate feedback.
Candidate 10 is the inner sweep's V1-lam0.1 variant at outer scale: candidate
7's k-NN conditional-rank-bootstrap machinery, candidate 9's zero-anchor
participation regime with its two poisons removed (the re-entry-pool
restriction dropped, Q0 targets held memory-exempt), and a FIXED lambda = 0.1
donor-coordinate blend for the non-Q0 targets. No calibration stage: lambda is
fixed, not chosen against any score.

Frozen spec: issue #42 comment 4902561460. Scored under the amended gate
(gates.yaml gate_1, PR #57/#59) exactly as run 11: runs-view c2st demoted,
gated benefit-space block folded into the geometry verdict, >=4/5 geometry AND
>=4/5 battery AND pooled Q0 (abs <= 5).

Verdict, exactly as computed: gate_1_pass = FALSE -- geometry 3/5 (needs 4/5),
battery 4/5 (passes), pooled Q0 +0.038% (passes). Seeds 0 and 1 each clip the
pairs-view C2ST (0.5315, 0.5330 vs 0.53); seeds 2/3/4 clear it. The only
battery failure is seed 0's 2-year autocorrelation (0.7896, deviation 0.0595 >
0.05) -- a short-lag overshoot from the memory injection, not the 10-year
undershoot of earlier candidates (every seed clears 4yr and 10yr).

Inner-vs-outer: the ~0.005-0.01 heat correction the inner sweep predicted for
the pairs C2ST delivered a reshuffle, not a uniform cooling -- mean cooled
-0.0014 (inner 0.5293 -> outer 0.5279), but seeds 0/1 heated up (crossing to
fail) while the worst inner seed (3, 0.5365) cooled to a comfortable pass
(0.5260). Still 3/5, one seed short -- exactly the forecast's named failure
(comment 4902561584, P(pass) ~0.42). The Q0 program was solved: pooled Q0
generalized inner +1.19% -> outer +0.04%, and the generated-vs-real all-zero
share gap is +0.0007 (no never-worker resurrection).

Generation verified byte-identical to the inner sweep's generate_variant at
memory_mode=lambda_blend, lam=0.1, use_zero_anchor_gate=True under candidate-7
seeding. Seed-0 reproduction run live in .venv-gate (float-exact). Full pytest
green in the repo .venv (361 passed, 22 skipped). ruff + black -l 79 clean.

Deliverables:
- scripts/run_gate1_candidate10.py -- deterministic; candidate-7 two-element
  substreams; the registered rules exactly.
- runs/gate1_rank_knn_v4.json -- schema gate1_rank_knn.v4; spec_registration =
  the candidate-10 comment; standard diagnostics + Q0 participation numbers +
  per-seed benefit_space + the amended scorecard.
- tests/test_gate1_qrf_candidate10.py -- seed-0 reproduction (skipif PSID +
  importorskip populace.fit) + generation-equivalence to the inner sweep + the
  Q0-exempt/full-pool checks + the amended-gate consistency block.

Links: issue #42; PRs #63 (inner-validation harness + design sweep), #62
(candidate 9), #61 (Q0 forensics), #55 (candidate 7).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant