Skip to content

Refit distributional tail and validate state factors - #13

Merged
MaxGhenis merged 1 commit into
mainfrom
model-fixes
Aug 7, 2026
Merged

Refit distributional tail and validate state factors#13
MaxGhenis merged 1 commit into
mainfrom
model-fixes

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Sol model-fix round per the adversarial statistical review, gated by independent byte-identical reproduction of all committed artifacts:

  • Tail refit at the attachment depth: scale 0.4674 → 0.2713 (SE 0.0198), Pareto index 3.69 — finite variance with margin; physical cap on draws (0.11% of weighted draw mass removed); mean-excess-by-cutoff diagnostics in the JSON
  • State factors shipped and validated on dollars: factor-adjusted 53-state dollar-rate MAE 0.928pp equal-state / 0.834pp issuance-weighted (corr 0.90); per-state level-gap table published
  • Shipped configuration (bootstrap + redraw) validated seed-averaged; seven jurisdictions remain outside the [0.7, 1.4] level gate even factored (AK, HI, ID, MN, SD, VI, WY) — the app model mode remains disabled; honest finding: the frozen configuration worsens raw coverage (7 of 9 cells beyond 3pp, all negative)
  • 74 tests; two byte-identical regenerations plus an independent third

🤖 Generated with Claude Code

@MaxGhenis
MaxGhenis merged commit 60c0ec2 into main Aug 7, 2026
1 check passed
@MaxGhenis
MaxGhenis deleted the model-fixes branch August 7, 2026 18:15
MaxGhenis added a commit that referenced this pull request Aug 7, 2026
Three adversarial reviews of revision 3 (methodology, red-team,
round-diff; archived verbatim in paper/reviews/round-3/) found that the
revision fixed one instance of a three-instance staleness event and
mis-attributed the coverage change. All findings applied:

The validation paragraph now names the true coverage channel, verified
in the PR #13 diff: the evaluation moved from the through-FY2023
primary model to the shipped frozen-through-FY2022 configuration
(62,984 training deviators against 79,919); none of the three fixes
enters the coverage computation. The two stale siblings from the same
refit are corrected with supersession notes: sign AUC 0.700 → 0.686
(E4) and E6's pre-#13 NY/CO comparison → the current ratio form
(NY 0.508, CO 0.692), which conveys a larger raw gap. "Structural"
weakens to "not specific to the gradient-boosted quantile stack" with
shared-cause candidates named. The sign-disagreement passage states the
robust interval-exclusion form (all three CIs exclude +$7.1M, maximum
upper bound +$0.6M), imports the artifact's own warning against reading
the model's sign as causal, grounds lever retirement in the
identity-versus-prediction mismatch, and names the classifier leg. The
review-history count is consistent everywhere (three rounds, seventeen
reports); the raw-facts leg's 845/856 agreement is disclosed beside the
certified chain's 856/856.

Attestations move into the repo: the five engine-leg manifests
(hash-verified against the pins in COUNTERFACTUAL.md) under
paper/snapshot/engine-leg/ with a provenance README, and a QRF
cross-process determinism attestation. scripts_build_data.py accepts
the documented arguments (output byte-identical). FACTS B3/E5/E9/I5
residue corrected. app/public/paper/ is re-rendered from this source,
resolving the stale published copy.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Aug 7, 2026
…ision 3) (#18)

* Align paper, docs, and pipeline provenance with the model-primary release

The manuscript's validation section described the pre-correction
pipeline: coverage within 3 points at seven of nine levels and the
model mode disabled pending fixes. The corrected pipeline (PR #13)
attains within-3-points coverage at only two of nine levels, and the
fixes shipped. Revision 3 states the current numbers with a disclosure
of the change, adds the pre-registered QRF benchmark as evidence the
under-coverage is not specific to the shipped estimator, reports the
Colorado engine leg end to end — the SMD-in-the-editing circularity,
the censoring-bounded repricing, and the sign disagreement between the
model-implied and accounting-reverse cost-share changes — and updates
the limitations and review-history passages to the deployed reality.

FACTS.md corrects E3 with a supersession note, updates E6 and the
tool-status prohibition, and adds section I tracing every claim of the
model-primary round to its committed artifact. The pipeline's
browser_consumer_status string and FINDINGS narrative now describe the
live consumer; regeneration is otherwise byte-identical (111 tests).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Apply round-3 review findings and re-render the published manuscript

Three adversarial reviews of revision 3 (methodology, red-team,
round-diff; archived verbatim in paper/reviews/round-3/) found that the
revision fixed one instance of a three-instance staleness event and
mis-attributed the coverage change. All findings applied:

The validation paragraph now names the true coverage channel, verified
in the PR #13 diff: the evaluation moved from the through-FY2023
primary model to the shipped frozen-through-FY2022 configuration
(62,984 training deviators against 79,919); none of the three fixes
enters the coverage computation. The two stale siblings from the same
refit are corrected with supersession notes: sign AUC 0.700 → 0.686
(E4) and E6's pre-#13 NY/CO comparison → the current ratio form
(NY 0.508, CO 0.692), which conveys a larger raw gap. "Structural"
weakens to "not specific to the gradient-boosted quantile stack" with
shared-cause candidates named. The sign-disagreement passage states the
robust interval-exclusion form (all three CIs exclude +$7.1M, maximum
upper bound +$0.6M), imports the artifact's own warning against reading
the model's sign as causal, grounds lever retirement in the
identity-versus-prediction mismatch, and names the classifier leg. The
review-history count is consistent everywhere (three rounds, seventeen
reports); the raw-facts leg's 845/856 agreement is disclosed beside the
certified chain's 856/856.

Attestations move into the repo: the five engine-leg manifests
(hash-verified against the pins in COUNTERFACTUAL.md) under
paper/snapshot/engine-leg/ with a provenance README, and a QRF
cross-process determinism attestation. scripts_build_data.py accepts
the documented arguments (output byte-identical). FACTS B3/E5/E9/I5
residue corrected. app/public/paper/ is re-rendered from this source,
resolving the stale published copy.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant