Refit distributional tail and validate state factors - #13
Merged
Conversation
MaxGhenis
added a commit
that referenced
this pull request
Aug 7, 2026
Three adversarial reviews of revision 3 (methodology, red-team, round-diff; archived verbatim in paper/reviews/round-3/) found that the revision fixed one instance of a three-instance staleness event and mis-attributed the coverage change. All findings applied: The validation paragraph now names the true coverage channel, verified in the PR #13 diff: the evaluation moved from the through-FY2023 primary model to the shipped frozen-through-FY2022 configuration (62,984 training deviators against 79,919); none of the three fixes enters the coverage computation. The two stale siblings from the same refit are corrected with supersession notes: sign AUC 0.700 → 0.686 (E4) and E6's pre-#13 NY/CO comparison → the current ratio form (NY 0.508, CO 0.692), which conveys a larger raw gap. "Structural" weakens to "not specific to the gradient-boosted quantile stack" with shared-cause candidates named. The sign-disagreement passage states the robust interval-exclusion form (all three CIs exclude +$7.1M, maximum upper bound +$0.6M), imports the artifact's own warning against reading the model's sign as causal, grounds lever retirement in the identity-versus-prediction mismatch, and names the classifier leg. The review-history count is consistent everywhere (three rounds, seventeen reports); the raw-facts leg's 845/856 agreement is disclosed beside the certified chain's 856/856. Attestations move into the repo: the five engine-leg manifests (hash-verified against the pins in COUNTERFACTUAL.md) under paper/snapshot/engine-leg/ with a provenance README, and a QRF cross-process determinism attestation. scripts_build_data.py accepts the documented arguments (output byte-identical). FACTS B3/E5/E9/I5 residue corrected. app/public/paper/ is re-rendered from this source, resolving the stale published copy. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MaxGhenis
added a commit
that referenced
this pull request
Aug 7, 2026
…ision 3) (#18) * Align paper, docs, and pipeline provenance with the model-primary release The manuscript's validation section described the pre-correction pipeline: coverage within 3 points at seven of nine levels and the model mode disabled pending fixes. The corrected pipeline (PR #13) attains within-3-points coverage at only two of nine levels, and the fixes shipped. Revision 3 states the current numbers with a disclosure of the change, adds the pre-registered QRF benchmark as evidence the under-coverage is not specific to the shipped estimator, reports the Colorado engine leg end to end — the SMD-in-the-editing circularity, the censoring-bounded repricing, and the sign disagreement between the model-implied and accounting-reverse cost-share changes — and updates the limitations and review-history passages to the deployed reality. FACTS.md corrects E3 with a supersession note, updates E6 and the tool-status prohibition, and adds section I tracing every claim of the model-primary round to its committed artifact. The pipeline's browser_consumer_status string and FINDINGS narrative now describe the live consumer; regeneration is otherwise byte-identical (111 tests). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Apply round-3 review findings and re-render the published manuscript Three adversarial reviews of revision 3 (methodology, red-team, round-diff; archived verbatim in paper/reviews/round-3/) found that the revision fixed one instance of a three-instance staleness event and mis-attributed the coverage change. All findings applied: The validation paragraph now names the true coverage channel, verified in the PR #13 diff: the evaluation moved from the through-FY2023 primary model to the shipped frozen-through-FY2022 configuration (62,984 training deviators against 79,919); none of the three fixes enters the coverage computation. The two stale siblings from the same refit are corrected with supersession notes: sign AUC 0.700 → 0.686 (E4) and E6's pre-#13 NY/CO comparison → the current ratio form (NY 0.508, CO 0.692), which conveys a larger raw gap. "Structural" weakens to "not specific to the gradient-boosted quantile stack" with shared-cause candidates named. The sign-disagreement passage states the robust interval-exclusion form (all three CIs exclude +$7.1M, maximum upper bound +$0.6M), imports the artifact's own warning against reading the model's sign as causal, grounds lever retirement in the identity-versus-prediction mismatch, and names the classifier leg. The review-history count is consistent everywhere (three rounds, seventeen reports); the raw-facts leg's 845/856 agreement is disclosed beside the certified chain's 856/856. Attestations move into the repo: the five engine-leg manifests (hash-verified against the pins in COUNTERFACTUAL.md) under paper/snapshot/engine-leg/ with a provenance README, and a QRF cross-process determinism attestation. scripts_build_data.py accepts the documented arguments (output byte-identical). FACTS B3/E5/E9/I5 residue corrected. app/public/paper/ is re-rendered from this source, resolving the stale published copy. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Sol model-fix round per the adversarial statistical review, gated by independent byte-identical reproduction of all committed artifacts:
🤖 Generated with Claude Code