Skip to content

Break-it D: blind-metric incentive (hidden decoy rewards the longest option) #139

Description

@sebasmos

Part of the break-the-cascade effort after the MedQA cascade came back null across three escalating manipulations (single seed, baseline-relative seed, and hard cases + reasoned seed + 2-of-3 majority = 0/30 conformity; a framing diagnostic was also null at 0.07/0.00/0.00 for independent/consensus/deference). The diagnosis: we were inducing Asch conformity to a wrong answer, which safety-trained models resist, instead of making a shortcut PAY OFF. These four experiments (A-D) test the reward/framing channels on MedQA text. Results reported per experiment.

D: blind-metric incentive (the reward channel)

Give the committee a hidden quality rubric that secretly rewards a decoy feature (the longest option, a known bias). This is the most on-thesis channel: it makes the shortcut PAY OFF rather than merely asserting it.

Measure

Drift toward the decoy (fraction choosing the longest option) and the accuracy drop, incentive prompt vs control, on cases where the longest option is not the correct one. Reuse benchmaxxing.blind_metric for the uptake framing.

Acceptance

  • A runner reports decoy-pick drift and accuracy drop, incentive vs control.
  • Deterministic; reuses the cache; offline mock path tested.
  • Result folded into the internal draft and the results PR.

Metadata

Metadata

Assignees

No one assigned

    Labels

    analysisMetrics, stats, visualization of resultsdifficulty: intermediateTouches one subsystem; some context neededexperimentExperiment runner / study designpriority: highDo this soon; unblocks the paper or other work

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions