Skip to content

v0.4.1 — eval refinements (+15pp validated)

Choose a tag to compare

@alexei-led alexei-led released this 24 Jun 18:26
· 10 commits to main since this release
266bcda

Patch: eval/test refinements. SKILL.md is unchanged from v0.4.0.

Changes

  • The eval shim auto-loads a local .env (gitignored), so runs are self-contained.
  • Fixed unfair assertions in the treat-urgent-auth eval case. Its prompt gives no code, so the disciplined model correctly asks for it first — the old assertions demanded it "verify the fix," which was unsatisfiable. Now: treat auth as high-risk, ask before blind-patching, commit to verify/check-regressions before declaring done, no fake success.

Eval result (rerun, baseline mode, gpt-5.4-mini)

  • with-skill 85% vs without-skill 70% → +15pp (per-assertion pass rate).
  • The treat-urgent-auth case now passes 4/4 with-skill vs 2/4 without — the fix is validated and the skill shows clear lift there.
  • Caveat unchanged: N=1 per case, so this is directional. Run eval/run-skill-evals.sh with more iterations for steadier numbers.