Skip to content

skill-placebo v0.1.0 (v1 results)

Choose a tag to compare

@simonether simonether released this 06 Oct 00:29
· 2 commits to main since this release

Superseded by v0.1.1: pass rates corrected (METHOD.md amendment 19); the agents' logs stay attached here.

skill-placebo v0.1.0: first results (v1) of a placebo-controlled test of the most-starred coding-agent skills.

2 of 9 skills beat a same-length placebo, both at Holm-adjusted p = 0.049 (0 worse, 7 no better). Secondary, against no skill: none of the 9 skills was measurably cheaper than running without a skill; a same-length placebo alone changed cost by +2% to +16%.

Claude Code (claude-opus-5-5), 15 public tasks, 30 trials per arm, 450 trials; the Codex numbers are the
pilot's kill test. Method registered before the first run, with dated amendments: METHOD.md.

Assets, the agents' full logs (raw stream and ATIF trajectory per trial, scrubbed), under the same paths as results/:

  • skill-placebo-v1-claude-code-main-agent-logs.tar.gz (sha256 306ff51dff39426ec88ab0e2197ac7b6834769b2bb25ca952683cdf1b793443b)
  • skill-placebo-v1-codex-pilot-agent-logs.tar.gz (sha256 1fd55e264728cf603a711a64f913aa04e442b66dbef35538d955bbcb2b753300)

Install: uvx skill-placebo run DietrichGebert/ponytail