Skip to content

Complete the Claude thinking supplement: Fable 5 and Sonnet 5 runs - #141

Merged
MaxGhenis merged 1 commit into
mainfrom
thinking-supplement
Aug 8, 2026
Merged

Complete the Claude thinking supplement: Fable 5 and Sonnet 5 runs#141
MaxGhenis merged 1 commit into
mainfrom
thinking-supplement

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Summary

🤖 Generated with Claude Code

Adds the POLICYBENCH_CHUNK_OVERRIDE=none escape hatch (test-locked, default
untouched) so the chunked Claudes' sensitivity runs share the canonical
whole-scenario shape, and extends sensitivity/claude-thinking-2026-08.md
to all three thinking-by-default models. With tool_choice auto: Fable 5
scores 86.9 weighted exact versus 79.9 on the board (would rank second,
at $0.323/household versus its $0.541 chunked board run), Opus 5 85.6
versus 79.8 (third), and Sonnet 5 80.2 versus 69.4 (+10.8, the largest
delta, with 56 of 1,984 answers failing to parse under auto and scored
as misses). Fable and Sonnet deltas combine de-chunking with thinking;
Opus 5's isolates thinking. All three prediction sets are attached to
the dashboard-data-20260805 release. The leaderboard is unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@vercel

vercel Bot commented Aug 8, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
policybench-site Ready Ready Preview Aug 8, 2026 7:41pm

Request Review

@MaxGhenis
MaxGhenis merged commit 3adcf79 into main Aug 8, 2026
6 checks passed
@MaxGhenis
MaxGhenis deleted the thinking-supplement branch August 8, 2026 19:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant