Releases: B0yko/proof-of-done
Release list
proof-of-done 0.1.1
First release on PyPI. The CLI now installs and runs without the git+https form:
uvx proof-of-done audit --demo
uv tool install proof-of-doneThe Claude Code plugin install is unchanged:
claude plugin marketplace add B0yko/proof-of-done && claude plugin install proof-of-done@proof-of-doneChanged
- README redesign: banner, badges, highlights, claim-type and verdict-reason tables, an at-a-glance results table generated from the committed results, detailed tables in collapsible sections, and absolute links so the page also renders on PyPI.
- The terminal recording uses the same palette as the banner; a social preview image is added.
Fixed
- The hook-timeout contract test kills the whole process group, so it also holds where
sh -cforks instead of exec-ing (dash on Ubuntu).
No change to the hook's behaviour or to the evaluation results. Full history: CHANGELOG.md.
proof-of-done 0.1.0
First release of proof-of-done: a Claude Code plugin whose Stop and SubagentStop hooks block a turn that ends with "tests pass", "the build succeeds", "lint is clean", "fixed", "deployed" or "verified" unless the session's own transcript shows the matching command ran after the last relevant edit and succeeded. The same engine runs after the fact as proof-of-done audit, which reports unsupported-claim rates and exports sessions as agent-trace/v1. There is no LLM call anywhere in the product.
Install
claude plugin marketplace add B0yko/proof-of-done && claude plugin install proof-of-done@proof-of-doneTry the audit on the bundled synthetic demo corpus:
uvx --from git+https://github.com/B0yko/proof-of-done proof-of-done audit --demoRequires python3 ≥ 3.9 on Claude Code's PATH (macOS or Linux). Tested with Claude Code 2.1.281.
Results (details and reproduction commands in the README)
- Held-out set (112 individually authored sessions, frozen before the claim detector existed): claim detection precision 98.6%, recall 84.0%; gate precision 98.2%, recall 82.1% for unsupported claims; 1.0% of turns without an unsupported claim were blocked.
- Templated set: gate precision 100.0%, recall 96.3% (measures consistency with the specification, written by the same author as the detector).
- Adversarial set: 34 of 36 gaming attempts met their labelled outcome; the two misses (an echo-only Makefile
test:target and annpm run test:watchscript) are documented limitations. - Hook latency, warm config cache, p95 on a MacBook Air (Apple M5, 24 GB): 175 ms at 10 MB and 578 ms at 50 MB with the system
python33.9; 113 ms and 467 ms with CPython 3.12. Audit throughput 140 MB/s.
All data is synthetic and authored in the repository. See the README's Limitations section, in particular the detector's missed phrasings (type-check wording is the largest gap) and the scope of the tamper check.
Full list of changes: CHANGELOG.md.