Repository navigation
Releases: simonether/skill-placebo
Release list
skill-placebo v0.1.1 (pass-rate correction)
skill-placebo v0.1.1: a correction to the v1 pass rates, and packaging fixes.
Correction (METHOD.md, amendment 19). Two agent timeouts whose tests passed afterwards had been counted as passes;
METHOD.md sections 5 and 7 say a timeout is a failed trial, and they now count as failed. planning-with-files
moves from "no better than placebo" to "worse than placebo" (pass-rate difference −20 pp [−37, −7], Holm-adjusted
p = 0.043). Costs and every cost ratio R are unchanged. Headline: 2 of 9 skills beat a same-length placebo, both
at Holm-adjusted p = 0.049 (1 worse, 6 no better).
Also:
- README: skills run with the scripted approval turn are marked, as METHOD.md section 5.1 requires.
- PyPI: the project page links to the repository and the issue tracker, with keywords and classifiers.
skill-placebo listandrunnow use this release's tag (v0.1.1).
The per-trial records in results/ are the same as in v0.1.0; only how two of them are counted changed. The
agents' full logs stay attached to release v0.1.0:
https://github.com/simonether/skill-placebo/releases/tag/v0.1.0 (sha256 in results/*/*/AGENT_LOGS.json).
This release has no assets.
Install: uvx skill-placebo run DietrichGebert/ponytail
skill-placebo v0.1.0 (v1 results)
Superseded by v0.1.1: pass rates corrected (METHOD.md amendment 19); the agents' logs stay attached here.
skill-placebo v0.1.0: first results (v1) of a placebo-controlled test of the most-starred coding-agent skills.
2 of 9 skills beat a same-length placebo, both at Holm-adjusted p = 0.049 (0 worse, 7 no better). Secondary, against no skill: none of the 9 skills was measurably cheaper than running without a skill; a same-length placebo alone changed cost by +2% to +16%.
Claude Code (claude-opus-5-5), 15 public tasks, 30 trials per arm, 450 trials; the Codex numbers are the
pilot's kill test. Method registered before the first run, with dated amendments: METHOD.md.
Assets, the agents' full logs (raw stream and ATIF trajectory per trial, scrubbed), under the same paths as results/:
- skill-placebo-v1-claude-code-main-agent-logs.tar.gz (sha256 306ff51dff39426ec88ab0e2197ac7b6834769b2bb25ca952683cdf1b793443b)
- skill-placebo-v1-codex-pilot-agent-logs.tar.gz (sha256 1fd55e264728cf603a711a64f913aa04e442b66dbef35538d955bbcb2b753300)
Install: uvx skill-placebo run DietrichGebert/ponytail