v0.4.2 — theory & evidence in README
Docs-only patch. SKILL.md and the eval are unchanged from v0.4.1.
Adds a "How it works — theory and evidence" section to the README, validated against the literature (via perplexity) with each claim labelled by evidence strength:
- Mechanism — procedures work via in-context conditioning; induction/iteration heads and CoT as a plausible substrate (Wei 2022, Olsson 2022, ICML/NeurIPS 2024). Grounded substrate, not a complete theory.
- Procedures, not personas — a supported guideline (Zheng EMNLP 2024), not a universal law.
- Self-selection — holds for strong/calibrated models (Route-to-Reason), not categorically better than a router.
- What the checklists encode — highest-evidence levers (self-debug, self-correction degrades, context-before-editing, verifier gaming); no clean RCT, hence our own eval.
- Our eval — +15pp, one run/one model: directional, not a settled effect.
- Explicitly debunks the unsourced "~150–200 instructions" figure.