0.2.0
LLM-as-judge scoring.
derail judge <report.json>: scores saved A/B reports on three rubric
constructs — belief_stickiness (contradiction probes), catastrophizing
and negativity (task turns) — with any OpenAI-compatible judge model- The judge sees the planted stimulus and the response text only, runs at
temperature 0; parse failures are counted, never silently dropped - Self-judging produces a warning;
--judge-model scriptedis an offline
dry-run - Zero new dependencies: the judge rides the existing OpenAI-compatible
client
Emulation, not diagnosis.