Apply the eight-reviewer manuscript audit - #124
Merged
Conversation
Eight parallel reviewers read the paper section by section and verified ~290 quantitative claims against the frozen snapshot. The fixes, most severe first: - The weighting mechanism was misstated: zero-reference rows do not contribute zero weight (they hold 66.9% of weight mass — which is exactly why the always-zero baseline scores 66.9%). The corrected account — weight concentrates in high-dollar, rarely-zero output groups — now lives once in the headline-metric section, with the abstract, related work, and baselines shrunk to references. - Wide tables (provenance, model runs, scope, deviation audit, parse audit) clipped mid-word at the page margin; they now render as pandoc pipe tables that wrap, and the leaderboard drops its constant Parsed/Total columns that overflowed. - The abstract's "newest flagship is not the strongest" was false on this board (Claude Fable 5 is Anthropic's newest and strongest); reframed as the within-family Opus 4.8 < 4.7 observation. - Stale §4.3 prose said GPT-5.5 ran pinned-low with a default rerun "scheduled," contradicting its own table; the frozen wave is the default-effort run and the prose now says so. - Claude Fable 5's cost/latency rendered as $0.000/0 s from empty per-row usage; the cost table now reconciles against modelStats ($0.541/98 s — the priciest model, not the cheapest), treats sub-second batch dispatch times as unmeasured (—), and caveats latency comparability across serving regimes. - "Joint hit rate is consistently lower than either marginal" failed for the top model; the Limitations bullet contradicted Appendix A on where failed attempts live; both corrected. Reference-defect terminology unified as "reference-computation defects." - The stale June-era annotations orphan inside the snapshot dir (3,300 rows failing the manifest's own hashes) is removed; a reproducibility note documents the frozen scenarios' stale source_dataset label. Front and back matter are unnumbered; the figure is now cited; the editorializing deterministic-engine aside is cut; plus a dozen smaller precision and repetition fixes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Section-by-section multi-agent review of the manuscript before arXiv/SSRN submission: eight parallel reviewers (six section readers, a numbers auditor over ~290 quantitative claims, and a whole-paper editor), every finding verified computationally against the frozen snapshot before fixing.
Blockers fixed
Majors fixed
source_datasetlabel.Minors (~15)
Dangling "Impact weighting" cross-reference → anchored section refs; Table 11 caption falsified by its own columns → accurate rationale; formal indicator
1{pred = ref}→1{|pred − ref| ≤ 1}; bounded-score zero-branch strictness caveated; within-5% promise dropped; figure now cited; front/back matter unnumbered; ACA expanded; 18-group reconciliation explained; the two unrelated 84% figures disambiguated; "31 of 35 (89%)"; run-on splits; % formatting.Verification
~/Downloads/policybench-arxiv-source.tar.gz).🤖 Generated with Claude Code