Repository navigation
Two plugins join the collection, and the others learn what their own measurements showed. aspiration-self-monitoring asks the agent to review a result the way it will be used and to compare it with a source it did not write; hygiene-self-monitoring keeps a change to what the request names. handoff gains waiting for a turn that ends while its result is still out, and progress lets several sessions share one ledger. The measurement is stricter too: a judge panel and blind labels, a reader test for handoff, a written protocol for what a confirmatory claim needs, and published numbers labelled for what they are.
| Plugin | Version |
|---|---|
| executive-self-monitoring | 1.8.0 |
| epistemic-self-monitoring | 0.3.5 |
| persistence-self-monitoring | 0.3.5 |
| termination-self-monitoring | 0.3.3 |
| coverage-self-monitoring | 0.5.1 |
| handoff-self-monitoring | 0.5.1 |
| progress-self-monitoring | 0.7.0 |
| integrity-self-monitoring | 0.1.2 |
| aspiration-self-monitoring | 0.1.0 |
| hygiene-self-monitoring | 0.1.0 |
Added
- aspiration-self-monitoring 0.1.0: review before you close. The hook records which files the turn edited and whether each was reviewed after its last change, in the medium it is used in: a document is read, an image is viewed, code is run or rendered. A close with an edited file nobody reviewed is named. The
[ASPIRATION CHECK]compares the result, point by point, with a source the agent did not write (the original, the specification, the real input, the owner's words): Criterion, Reviewed, Found, Remainder (meets | defect | unverified | blocked). A close short of the objective is asked where the agent has not looked, with the count of files the session read against the files the project has. The skill is put in context at session start. Non-blocking by default;ASPMON_STRICT=1blocks once. Measured on Claude Code (Sonnet 5, strict): on the flag of Nepal, without the plugin 0 of 3 passed and none rendered; with the skill and the question 1 of 6 passed and 4 of 6 rendered. On the report and the guide cases the question changes nothing: report 3/3, guide 2/3, the guide's one failure the same falsemeetsas before. - hygiene-self-monitoring 0.1.0: keep a change to what the request names. From an outside proposal, narrowed by the bench: the failure that reproduced is a small change that reaches public behaviour the request did not name. The hook records the public functions a change reaches, file by file, and the close names what the request does not and puts it back or to the owner. Cursor: 6/6 with the plugin against 0/6 unaided on the header case. Claude Code, one run per cell: 4 unaided cells with scope failures, 0 with it.
HYGMON_STRICT=1blocks once. - handoff 0.5.0:
waiting, the close of a turn whose result is still out: a command in the background, a subagent, a job whose result decides what comes next.Waiting-onnames it and what happens with its result;Nextsays nothing until it ends. Two evals and aclose.status_notrule in the shared grader. - progress 0.7.0: several sessions on one ledger. A session that takes an open item leaves a claim with when it took it and until when it holds; a claim that runs out is a session that could not go on, and another takes the item over. The ledger is the repository's: from a linked git worktree it is the main checkout's
.agent/progress.md. No lock: the protocol makes a collision safe rather than impossible. - The handoff block judged as a summary for deciding (
bench/studies/handoff-fidelity/) and a reader test for handoff (bench/studies/handoff-reader-test/): handoff's outcome is in the reader, so a reader with no tools gets only the last message and answers what to do next. - docs/STUDY-PROTOCOL.md: what a confirmatory claim needs. Exploratory and confirmatory runs kept apart; a pre-registration committed before the first run; blind labels and a judge panel (
scripts/blind-labels.js,blind-sheet.js,claude-judge.js,cursor-judge.js) as the first pieces. - Skill variants in the runners (
--skill-variant <name>inclaude-bench.js,claude-eval.jsandcursor-bench.js), and new bench exercises for the outside proposals on aspiration and on a change-budget plugin. - The Claude runner refuses a shell command by name and lets the agent run node the way Claude writes it (
cd "<workspace>" && node ...), and reaches the web where a case needs it.
Changed
- handoff 0.5.1: under
waiting, Next says until when. A reader who took a bareNext: nothingto mean nothing is still to come was the finding of the triage reading (6 of 42 readings). - The behavioural eval grants what the agent composes commands with, and keeps every run's stream. A compound command under dontAsk needs every part granted.
- The judge-panel pilot found defects in its cases before anyone labelled it, and the cases were rebuilt.
- The published numbers are labelled for what they are. Cases selected on the outcome and not pre-registered, graders written after the runs: the numbers are exploratory, and the docs say so.
- The eval no longer reports a delta it cannot have. On a block case graded on the plugin's own marker, the arm without the plugin cannot write the marker.
- Related work on proxies and reward hacking added to
docs/RESEARCH.mdand to integrity's references: five references, each checked against its source. - Docs swept for the two new plugins: six plugins put their skill in context, five have an opt-in strict gate, ten plugins in the counts;
docs/DIRECTORY.mdstate at 2026-10-04 (0 blocking findings,claude plugin validatepasses on each plugin and on the marketplace).
Fixed
- The directory audit flags the web reached through the shell (
curl,wget,Invoke-WebRequest). - The drawing cases no longer pass a generated drawing kept in
tools/. - coverage 0.5.1: an
[ASPIRATION CHECK]is not read as deferred work. ItsFoundline describes what the review found ("still needs one more dash") and read as a deferral.
Full Changelog: v0.16.0...v0.17.0