0.9.0 — a record is written by a machine and trusted like one, and now something reads it
LatestFour work items, six issues. One sentence in four sets of clothes: a record is written by a machine and then trusted like one, while nothing checks what it says.
What ships
A loaded file may no longer name a version at or above the running one (#179, #98). The check read one number — the version in plugin.json — so a version written ahead of its release was green every day until the day it shipped, and red on that release's own preparation commit, hours in, after the broad gate had already run. Alongside it, three places credited -z alone with turning git's escaping of non-ASCII paths off; re-measured on git 2.50.1 over four quoting variants plus a control, either argument does it by itself.
The round record carries the reviewer's paste-ready fix, and a | inside a cell no longer truncates the row (#187, #189). The review skill requires a paste-ready fix for every 🔴 and every 🟡 and spends four paragraphs on what makes one paste-ready. round_record.py new copied the report's four tables and dropped everything else — so not one of those blocks reached the file a fix pass is told to open instead of the report.
evidence-check reads what a record says about the tree (#190). A ledger row is a claim about the tree that something reads; a spec.md, a plan.md, a round record or a phase record states the same kind of thing and nothing read it. A record of a work item that has not shipped is now refused when it names a backticked compound identifier nothing outside the records carries, and a path#unit@hash stamp in a record is resolved by the ledger's own reader. The boundary is a file the release already removes — the work item's ledger fragment — so there is no new state to maintain and no list to keep.
round_record.py new says what bound the next round is under, as it writes the record (#207). The chain bounds a run one step earlier than the cap: after a record whose floor row reads no, at most one later record may close on a fix, and the record that reads its fixes ends the run whatever it finds. That was enforced only at the broad gate — after every round had already been spawned. #179's own chain overran it by three rounds, and the gate refused 37.9 minutes and 180 calls of agent work that were then reverted.
What the release measured about itself
The #190 · #207 run was capped, and the cap did its job. Three review rounds and two fix passes. Round 2 reopened; round 3 read the fixes and ended the run whatever it found. It found seven more things and not one of them blocks — they were filed rather than fixed, which is exactly what docs/review-chain-spec.md §The reopening exists to produce: #217 through #222, beside #215 and #216 for the two answered questions.
Twice, a fix closed the coordinate and left the class one step over. Round 1's 🔴 2 became round 2's 🟡 1 became round 3's 🟡 2 — the printed bound disagreeing with the gate it exists to predict, in three different places, each time inside the fix for the one before. Round 3 settled it by construction rather than by reading: a differential of bound_line against chain_check.stopping_floor over all 584 record sequences of length ≤ 3, which found exactly one disagreeing class, 16 sequences, all permissive.
The Windows leg had been red since the commit that added the records arm — through three review rounds and two broad gates. A coordinate the arm built printed with the platform separator, so the same file read seal\specs\… from a record and seal/specs/… from a ledger row naming it. Every round and every gate ran on macOS, where the fix is a no-op. That is agent-contract §13 in its plainest form — a defence resting on a platform guarantee nobody removed. The case passes ntpath rather than skipping off Windows.
One rider had to be re-stamped at the release, and that is the squash's own cost. The stamp named a feature-branch commit the squash into the release branch discarded, so the rider-reachability check went red at the release rather than on the branch that wrote it — the class docs/branch-and-release.md records for release → main, arriving one direction earlier.
Every segment was measured at its boundary and posted to the flow log: 413 tool calls over roughly 100 minutes across four measured segments. Two implementer segments at 1.00–1.01 tools per turn against two reviewer segments never flagged for batching; the reviewers cost less than half of either fix pass and found eleven things the fix passes had not. The log also carries an instrument correction: the output-token undercount is not a scale factor — 3.2× on one segment and 334× on another, the same day — so no published reading involving a reviewer's output tokens is usable until that ships.
The gate
./bin/test 2541 passed, 2 skipped · uvx ruff check . and uvx ruff format --check . clean over 110 files · ./bin/evidence-check . exit 0 with 764 ok · 0 drifted · 0 broken · 0 external · 0 old-format · gather_changelog.py --check and fold_ledger.py --check both exit 0, no ledger fragment left and no open evidence-todo row. Every feature pull request was green on all six legs at the commit that merged it, Windows included.
Full changelog: CHANGELOG.md · Compare: v0.8.3...v0.9.0