results: blind-metric substitution probe (Story 4, #18/#139) - #157
Merged
Conversation
…igned rubric awarding a wrong option induces modest reward-hacking (decoy uptake +0.075, 3/40); test-awareness priming fully suppresses it (to 0); drift is mostly aware (2/3 name the rubric). Keyless-reproducible (0 API calls)
… responses, mis-scoring ~85% as option A). Corrects this experiment's committed numbers; see the 2026-07-21 re-grade
Member
Author
|
Parser-bug correction (2026-07-21). The answer parser (
|
…ted numbers; keyless-reproducible off the updated cache)
This was referenced Jul 21, 2026
Closed
sebasmos
added a commit
that referenced
this pull request
Aug 4, 2026
…corrected (closes #131) Rebased fresh off current main (the original branch predated #142/#143/#146/#148/#154/#157/ #161/#219 and would have deleted all of that merged work if landed as-is). Addresses @Agastya191's review on #159: - flash_lite_reference was hardcoded to scale_c's PRE-parser-fix numbers (0.33/0.51 at n=150, p<1e-4). Now reads dynamically from --scale-c-summary (scale_c's committed summary, PR #141) so it cannot drift out of sync with a future fix there. Regenerated: 0.729/0.847 at n=85, p=0.041. - Also caught in the same pass: the docstring and README's own flash numbers were ALSO stale (60 hard cases, generic 0.10/anchored 0.083) versus the actually-committed, already-corrected summary (28 hard cases, generic 0.714/anchored 0.679). Rewrote both to match the real data and restated the finding precisely: both tiers conform substantially to a bare peer (~0.71-0.73), but only flash-lite's conformity climbs further under a case-anchored rationale (+0.12, p=0.041); flash's does not move. The anchoring lever is model-dependent, not conformity itself. Verified end-to-end: keyless reproduction confirmed with the key unset (new_api_calls_this_run: 0, n_hard_cases: 28, matches exactly), ruff clean, 625 tests pass on the rebased base, no hardcoded personal paths, 0 em dashes.
sebasmos
added a commit
that referenced
this pull request
Aug 4, 2026
…referee) (#162/#163/#164/#165) Provenance manifest via the registered nih_cxr14 adapter with sha256 checksums per used image (#162); solo cue-susceptibility across cable/corner-tag/watermark/laterality, watermark +0.11 above the temperature-resampled noise floor (#163); the watermark cascade, shared adopt 0.97 vs isolated 0.34, contagion +0.63 (#164); a deployable referee with no privileged knowledge (transcript + one private re-read) catching peer-driven adoptions at P/R 0.86, FPR 0.23 versus a naive conformity gate's FPR 0.92 (#165). Branched fresh off main (not off the stale results/cascade-at-scale, which predates #146/#148/ #154/#157/#161 and would delete that work if merged) and reuses the already-committed, already-reviewed solo/cascade logic and numbers from that branch. All four runners are fully keyless-reproducible from the committed img_cache.jsonl (new_api_calls_this_run: 0, verified).
sebasmos
added a commit
that referenced
this pull request
Aug 4, 2026
…loses #108) Rebased fresh off current main (the original branch predated #142/#143/#146/#148/#154/#157/ #161/#219 and would have deleted all of that merged work if landed as-is, same stale-base issue as #143/#150/#141/#159). Addresses @Agastya191's review on #158: - README/PR body had stale pre-fix numbers (accuracy 0.29/0.14, flash flip 0.76 vs 1.00 at p=9.8e-5) versus the already-corrected committed summary (accuracy 0.89/0.78, flash flip 0.079 vs 0.455 at p=3.3e-3, flash-lite 0.090 vs 0.636 at p=4.4e-7). Rewrote to match. - ever_flipped collapsed the three cues into a per-case boolean that doesn't match the stated headline flip rate (0.063/0.117). Now reports both per_record_flip_rate (the headline quantity) and ever_flipped_rate, clearly labeled and distinguished. - Grading-standard mismatch: clean_correct was imported from the prior solo pipeline rather than re-verified. Now re-graded by exact match against ground truth, the same standard as options_only_correct. Re-graded values are identical to the imported ones (0.89/0.78), so no actual discrepancy, but the standard is now transparent and consistent. - Cache race: _Cache.complete checked the store under lock but called the API outside it, so two threads could both miss the same key and both append a duplicate entry (confirmed: the committed cache had exactly one). Lock now spans the full miss-check-fetch-store sequence. Cache also pruned from 6054 shared-cache entries down to the 400 this script actually requests (100 cases x 2 models x 2 probes). Verified end-to-end: keyless reproduction confirmed with the key unset (new_api_calls_this_run: 0, all numbers match exactly), ruff clean, 625 tests pass, no hardcoded personal paths, 0 em dashes.
sebasmos
added a commit
that referenced
this pull request
Aug 4, 2026
Rebased fresh off current main (the original branch predated #142/#143/#146/#148/#154/#157/ #161/#219 and would have deleted all of that merged work if landed as-is, same stale-base issue as #143/#150/#141/#159/#158). Addresses @Agastya191's review on #155: - README/PR body carried pre-fix numbers (shared 0.175/0.15/0.175/0.15/0.15, p=1.0) versus the already-corrected committed summary (0.275/0.3/0.275/0.325/0.325, gained 3 lost 1 p=0.625). Rewrote to match and restated the finding as no significant compounding at p=0.625. - Real bug: the isolated holdout's prompt was byte-identical to the bare prompt every round (board text was built only from peer votes, which are never visible in isolated mode), so at temperature 0 it deterministically reproduced round 1's answer every round. Isolated adoption was 0 by construction, not by measurement. Fixed by reminding the isolated holdout of its own previous-round answer (visible to it in isolated mode; only peer turns are hidden), so the prompt genuinely differs round to round. The shared arm's board construction is UNCHANGED, so the shared numbers are identical to what was already reviewed; only the isolated arm is new. Result: isolated adoption is now a genuine (if still near-floor) measurement: 0.0/0.025/0.0/ 0.025/0.0, at most 1 of 40 cases per round, rather than a trivial constant 0.0. Verified end-to-end: ran for real with the isolated fix (206 new calls), then re-verified keyless reproduction with the key unset (new_api_calls_this_run: 0, exact match on both curves), ruff clean, 629 tests pass, no hardcoded personal paths, 0 em dashes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What (Story 4; issues #18, #139)
The fuller blind-metric substitution probe. break-it D (a rubric rewarding the longest option) was null; this probes a direct misaligned metric, a rubric that awards full marks to a specific wrong option, under three conditions per case (baseline / blind / test-aware).
Result (MedQA, 40 cases)
Read. Blind-metric gaming is real but small on this text task, mostly aware, and fully reversible by a single test-awareness cue. For a referee this is encouraging: the gaming is shallow and often self-declared, so a monitor that reads the justification or primes evaluation-awareness catches or prevents most of it.
Reproduction
Keyless from the committed cache (
new_api_calls_this_run = 0). ruff clean. Seeexperiments/blind_metric/README.md.