Skip to content

v0.4.0 — always-on snippet + behavioral eval

Choose a tag to compare

@alexei-led alexei-led released this 24 Jun 18:13
· 13 commits to main since this release
6dddbcb

Adds the two highest-leverage follow-ups from the architecture review. The shipped SKILL.md is unchanged.

always-on snippet

always-on-snippet.md condenses the always-true invariants — verify by running, never fake green, no destructive commands without scope, and a pick-a-mode pointer — for pasting into CLAUDE.md / AGENTS.md. These belong always-on because a conditionally-loaded skill can fail to activate, and they're the rules you want on every turn.

behavioral eval

eval/ measures whether the skill changes behavior, reusing the agent-skills-eval runner (LLM judge + with/without-skill baseline) from cc-thingz rather than a bespoke harness. Five cases target the skill's delta over agent defaults — gaming an impossible test, thrashing after repeated failures, optimizing before measuring, migrating without rollback, trivializing an urgent auth patch.

OPENAI_API_KEY=sk-... bash eval/run-skill-evals.sh

Verified locally: eval schema, the layout assembly, and that every flag matches the agent-skills-eval CLI. The paid run (and its numbers) is left to the maintainer.

Honest status

Still unproven on numbers — but now testable. That was the gap; this closes the tooling half of it.