Skip to content

Align paper, docs, and provenance with the model-primary release (revision 3) - #18

Merged
MaxGhenis merged 2 commits into
mainfrom
alignment-pass
Aug 7, 2026
Merged

Align paper, docs, and provenance with the model-primary release (revision 3)#18
MaxGhenis merged 2 commits into
mainfrom
alignment-pass

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Revision 3 of the working paper plus the docs/pipeline provenance alignment, taken through a third adversarial review round (three reviewers, reports archived verbatim in paper/reviews/round-3/) with all findings applied.

What the revision corrects

  • The published validation section described the pre-PR-Refit distributional tail and validate state factors #13 pipeline (coverage within 3pp at seven of nine levels; model mode disabled). The current claim: the shipped frozen-through-FY2022 configuration attains within-3pp coverage at only two of nine levels (gaps −0.2 to −7.4pp, all negative), with the superseded figures preserved in the fact catalog and the true channel named and code-verified — the evaluation frame moved to the shipped frozen configuration; none of the three Refit distributional tail and validate state factors #13 fixes enters the coverage computation. Two stale siblings from the same refit (sign AUC 0.700→0.686; E6's pre-Refit distributional tail and validate state factors #13 NY/CO comparison → current ratio form NY 0.508 / CO 0.692) are corrected with supersession notes.
  • The roadmap reports the Colorado engine leg end to end: the SMD-in-the-editing circularity (certified chain 856/856 without the state rule; raw-facts leg 845/856 with diagnosed divergences), the censoring-bounded repricing (−$1.46..$0/case-month over 46 records censored at $165), and the sign disagreement — model-implied −$1.55M..−$2.20M/yr vs the +$7.1M accounting-reverse bound, stated in its robust form (all three bootstrap intervals exclude +$7.1M; two of three span zero), with the artifact's own warning against causal readings and lever retirement grounded in the identity-versus-prediction mismatch.
  • The QRF benchmark enters as evidence the under-coverage is not specific to the gradient-boosted stack, and the deployment reality (SMD scenario as the only lever, seven level-gated jurisdictions disabled) replaces every "model mode is disabled" claim across the manuscript, docs, FINDINGS narrative, and the pipeline's provenance strings (regeneration byte-identical outside the two intended strings).
  • Attestations previously external land in-repo: five engine-leg manifests hash-verified against the pins in COUNTERFACTUAL.md (paper/snapshot/engine-leg/), and a QRF cross-process determinism attestation. scripts_build_data.py now accepts the README's documented arguments (output verified byte-identical).
  • app/public/paper/ re-rendered from this source — the served copy previously still carried the pre-fix text.

Verification: 111 tests + CI-exact ruff green (scope extended to scripts_build_data.py); independent run_all regeneration reproduces every artifact byte-for-byte outside the two intended provenance strings; every quantitative claim in the changed regions verified against committed artifacts by the round-3 reviewers, including an independent recomputation of the 46/$165, 53-flip, and 856 counts from the pinned raw CSV and a byte-comparison of the live deployment against the repo.

🤖 Generated with Claude Code

MaxGhenis and others added 2 commits August 7, 2026 14:16
…ease

The manuscript's validation section described the pre-correction
pipeline: coverage within 3 points at seven of nine levels and the
model mode disabled pending fixes. The corrected pipeline (PR #13)
attains within-3-points coverage at only two of nine levels, and the
fixes shipped. Revision 3 states the current numbers with a disclosure
of the change, adds the pre-registered QRF benchmark as evidence the
under-coverage is not specific to the shipped estimator, reports the
Colorado engine leg end to end — the SMD-in-the-editing circularity,
the censoring-bounded repricing, and the sign disagreement between the
model-implied and accounting-reverse cost-share changes — and updates
the limitations and review-history passages to the deployed reality.

FACTS.md corrects E3 with a supersession note, updates E6 and the
tool-status prohibition, and adds section I tracing every claim of the
model-primary round to its committed artifact. The pipeline's
browser_consumer_status string and FINDINGS narrative now describe the
live consumer; regeneration is otherwise byte-identical (111 tests).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three adversarial reviews of revision 3 (methodology, red-team,
round-diff; archived verbatim in paper/reviews/round-3/) found that the
revision fixed one instance of a three-instance staleness event and
mis-attributed the coverage change. All findings applied:

The validation paragraph now names the true coverage channel, verified
in the PR #13 diff: the evaluation moved from the through-FY2023
primary model to the shipped frozen-through-FY2022 configuration
(62,984 training deviators against 79,919); none of the three fixes
enters the coverage computation. The two stale siblings from the same
refit are corrected with supersession notes: sign AUC 0.700 → 0.686
(E4) and E6's pre-#13 NY/CO comparison → the current ratio form
(NY 0.508, CO 0.692), which conveys a larger raw gap. "Structural"
weakens to "not specific to the gradient-boosted quantile stack" with
shared-cause candidates named. The sign-disagreement passage states the
robust interval-exclusion form (all three CIs exclude +$7.1M, maximum
upper bound +$0.6M), imports the artifact's own warning against reading
the model's sign as causal, grounds lever retirement in the
identity-versus-prediction mismatch, and names the classifier leg. The
review-history count is consistent everywhere (three rounds, seventeen
reports); the raw-facts leg's 845/856 agreement is disclosed beside the
certified chain's 856/856.

Attestations move into the repo: the five engine-leg manifests
(hash-verified against the pins in COUNTERFACTUAL.md) under
paper/snapshot/engine-leg/ with a provenance README, and a QRF
cross-process determinism attestation. scripts_build_data.py accepts
the documented arguments (output byte-identical). FACTS B3/E5/E9/I5
residue corrected. app/public/paper/ is re-rendered from this source,
resolving the stale published copy.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@MaxGhenis
MaxGhenis merged commit 012da3a into main Aug 7, 2026
1 check passed
@MaxGhenis
MaxGhenis deleted the alignment-pass branch August 7, 2026 21:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant