Skip to content

v0.4.2 — theory & evidence in README

Choose a tag to compare

@alexei-led alexei-led released this 24 Jun 20:21
· 8 commits to main since this release
473a25a

Docs-only patch. SKILL.md and the eval are unchanged from v0.4.1.

Adds a "How it works — theory and evidence" section to the README, validated against the literature (via perplexity) with each claim labelled by evidence strength:

  • Mechanism — procedures work via in-context conditioning; induction/iteration heads and CoT as a plausible substrate (Wei 2022, Olsson 2022, ICML/NeurIPS 2024). Grounded substrate, not a complete theory.
  • Procedures, not personas — a supported guideline (Zheng EMNLP 2024), not a universal law.
  • Self-selection — holds for strong/calibrated models (Route-to-Reason), not categorically better than a router.
  • What the checklists encode — highest-evidence levers (self-debug, self-correction degrades, context-before-editing, verifier gaming); no clean RCT, hence our own eval.
  • Our eval — +15pp, one run/one model: directional, not a settled effect.
  • Explicitly debunks the unsourced "~150–200 instructions" figure.