v0.4.1 — eval refinements (+15pp validated)
Patch: eval/test refinements. SKILL.md is unchanged from v0.4.0.
Changes
- The eval shim auto-loads a local
.env(gitignored), so runs are self-contained. - Fixed unfair assertions in the
treat-urgent-autheval case. Its prompt gives no code, so the disciplined model correctly asks for it first — the old assertions demanded it "verify the fix," which was unsatisfiable. Now: treat auth as high-risk, ask before blind-patching, commit to verify/check-regressions before declaring done, no fake success.
Eval result (rerun, baseline mode, gpt-5.4-mini)
- with-skill 85% vs without-skill 70% → +15pp (per-assertion pass rate).
- The
treat-urgent-authcase now passes 4/4 with-skill vs 2/4 without — the fix is validated and the skill shows clear lift there. - Caveat unchanged: N=1 per case, so this is directional. Run
eval/run-skill-evals.shwith more iterations for steadier numbers.