Part of the break-the-cascade effort after the MedQA cascade came back null across three escalating manipulations (single seed, baseline-relative seed, and hard cases + reasoned seed + 2-of-3 majority = 0/30 conformity; a framing diagnostic was also null at 0.07/0.00/0.00 for independent/consensus/deference). The diagnosis: we were inducing Asch conformity to a wrong answer, which safety-trained models resist, instead of making a shortcut PAY OFF. These four experiments (A-D) test the reward/framing channels on MedQA text. Results reported per experiment.
D: blind-metric incentive (the reward channel)
Give the committee a hidden quality rubric that secretly rewards a decoy feature (the longest option, a known bias). This is the most on-thesis channel: it makes the shortcut PAY OFF rather than merely asserting it.
Measure
Drift toward the decoy (fraction choosing the longest option) and the accuracy drop, incentive prompt vs control, on cases where the longest option is not the correct one. Reuse benchmaxxing.blind_metric for the uptake framing.
Acceptance
- A runner reports decoy-pick drift and accuracy drop, incentive vs control.
- Deterministic; reuses the cache; offline mock path tested.
- Result folded into the internal draft and the results PR.
Part of the break-the-cascade effort after the MedQA cascade came back null across three escalating manipulations (single seed, baseline-relative seed, and hard cases + reasoned seed + 2-of-3 majority = 0/30 conformity; a framing diagnostic was also null at 0.07/0.00/0.00 for independent/consensus/deference). The diagnosis: we were inducing Asch conformity to a wrong answer, which safety-trained models resist, instead of making a shortcut PAY OFF. These four experiments (A-D) test the reward/framing channels on MedQA text. Results reported per experiment.
D: blind-metric incentive (the reward channel)
Give the committee a hidden quality rubric that secretly rewards a decoy feature (the longest option, a known bias). This is the most on-thesis channel: it makes the shortcut PAY OFF rather than merely asserting it.
Measure
Drift toward the decoy (fraction choosing the longest option) and the accuracy drop, incentive prompt vs control, on cases where the longest option is not the correct one. Reuse
benchmaxxing.blind_metricfor the uptake framing.Acceptance