feat(hooks): record what this branch put on the board, and whether it was refined - #399
Conversation
CLOUD-491 A live `plan-hold` did not stop a container restart, and nothing records that it was live — the hold ships with no sensor
Why CLOUD-451 landed Measured 2026-08-12, ~22:12 UTC. A hold was armed ( So the container went down with four live tracked tasks, one of them a full Reproduced 2026-08-12, ~23:45 UTC, in the session grooming this issue: a hold was armed, and the container was restarted roughly a minute later, killing it. Two independent instances now, and the second one left artifacts the first did not — see the measurements below. The finding is not "the hold is wrong". It is that the hold cannot be graded. A mechanism that occupies the container leaves no record of having done so, so from inside there is no way to distinguish the hold having been live and reclaimed anyway from the hold having already exited from a platform event no occupancy could defer. All of them produce the same observation — a fresh container — and the issue's acceptance is written in terms nobody can check after the fact. That is the sensor-without-a-gate shape inverted: a mechanism with no sensor. It is why an instance can be reported but not diagnosed, and why the honest first move is evidence rather than a bigger hammer. Measured 2026-08-12 ~23:30–23:46, and it changes the mechanism The first draft of this issue assumed one kind of container replacement. There are two, they are indistinguishable from inside without a sensor, and they have opposite consequences for any sensor built from a local file. 1. The 23:30 boot destroyed everything. A session was demonstrably alive at 23:28 — this issue's own last edit — and nothing it wrote survived anywhere writable. 2. The 23:45 boot preserved the disk. Same appearance from inside, opposite consequences. A heartbeat file under 3. The surviving hold directory was empty. 4. The last 182 seconds of writes did not survive, and this falsifies the predicate this issue shipped with. Every surviving pre-boot write stops at 23:42:15 — the MCP logs of four different servers, Serena's log, So the first draft's predicate — The heartbeat question, stated so it can be settled rather than argued. Two different things are called a heartbeat here and only one is cheap:
The first is a prerequisite for deciding the second. Ship the sensor, read it, then decide. The activity-versus-existence question — every wait in this repo that survives does I/O, and this one does not — is CLOUD-500, deliberately blocked on this issue for that reason. What the sensor must distinguish, and how The hold records why it stopped, not merely when it last ran — absence of an intentional-exit record is the signal, and absence is the one reading robust to losing the tail. Two line kinds in one appended file:
The residual error is bounded and in the conservative direction: a hold released inside the lost-write window reads as "was live". The residue probe of the first draft does not work either, and is replaced. It separated the last two rows by whether prior-container residue exists, naming Refinement — Ready
Acceptance
Measured in the session that also filed CLOUD-488; the plan whose approval button was destroyed was the one grooming CLOUD-427. The 23:30–23:46 measurements and the empty-hold-directory artifact were added in the grooming session, which was itself restarted twice while doing it. CLOUD-514 Nothing prices filing over fixing, so spinning off a defect in the PR's own diff is arithmetically cheaper than finishing it
Why Every gate in this repo prices failing to record something. Nothing anywhere prices the opposite: recording something instead of doing it. Filing satisfies every one of those gates at once and costs a few seconds, while finishing costs a diff, a suite and a landing. For an agent under pressure that is not a temptation, it is arithmetic — and the board becomes the escape hatch every guardrail points at. AGENTS.md already names the behaviour: "A punt is any deferral you could have closed … offering an action you are already authorized to take." That rule is prose, and prose is feedforward only. Nor is the substitution a fair trade. Across studies of admitted technical debt only 26.3–63.5% of it is ever removed, with median lifespans of 18–172 days and instances surviving more than ten years; in trackers specifically the repayment distribution is severely skewed, median 25 hours against a mean of 872 hours. A ~35× median/mean gap is the signature of a long tail never repaid at all. Filing does not defer a fix, it converts one into a weighted coin-flip. Measured 2026-08-13, PR #390. CLOUD-513 is a defect in code written in that PR: two new fixture suites read ambient git config, passed No reviewer is present at the moment of the choice, so the cost has to land on the author. Landing here is trunk-based: a branch fast-forwards onto Two mechanisms are ruled out before any is proposed 1. Judging the spin-off is forbidden. "Is this issue related enough to the PR to belong in it?" and "should this have been fixed instead?" are both model verdicts, which non-negotiable 3 refuses: a gate resolves to a command and an exit code over an object it decides. CLOUD-505 hit the identical wall, and its resolution is the template — do not judge the content, price the action. 2. A time window is measured, and rejected. The obvious credential-free proxy for "issues this branch filed" is "issues created between the PR opening and its merge". Measured over the last 40 merged PRs against every It fires on essentially everything, and structurally rather than tunably: this repo runs a fleet, so a window captures every session's filings. PR #373's window alone holds 20 issues, none of them its author's. A gate at 99.5% is switched off the first day. That measurement is what makes a receipt load-bearing rather than convenient. Which branch filed which row has never been recorded anywhere — not in git, not on the PR, not on the board. Without a record the predicate is not merely noisy, it is not computable at all. Refinement — Ready Refinement gate: Definition of Ready & Done. This body carries only specializations.
Chat is not a fourth sink, and that is a memory property rather than a style rule A finding stated in prose and never written durably is context-resident state. Long-horizon agents lose exactly that: unbounded context forces truncation and summarisation, attentional dilution degrades recall of any single item as context grows, and material in the middle of a long context is recalled least reliably — which is why an external store holding raw records at full fidelity is the standing recommendation. So a finding left in an aside is not merely undocumented, it is predictably lost, and the container reclaim this session survived is only the abrupt version of the same thing. How this gets satisfied without doing the work Three, found by attacking the mechanism rather than reasoning about it, each with local evidence. Two are closed above by changing the design; the third is a residual, stated with its failure direction. A ready-lint verdict is forgeable if the agent supplies the payload — closed by moving the lint into the recorder, which lints the tracker's response instead. The finding stands and is why §1 reads as it does; the original remedy (a second receipt, made trustworthy elsewhere) was the wrong fix for it. The measurement that made the hole real. An unfiltered listing launders sink 2 — closed by dropping the requirement, because it bought nothing. The record survives a branch restart — and that is fail-closed here, which is why it is a residual and not a blocker. What this cannot do It cannot judge the reason, and a minimal-compliance path still exists — a Ready block written to satisfy Gating creates pushes pressure toward commenting instead, and that is accepted rather than fixed. A comment on the row that already owns a finding is a legitimate durable home — cheaper than filing by design, since the friction is meant to sit on the impulsive path and not on honest recording. What makes it acceptable is that the pressure runs toward recording in the right place, not toward silence. It is still recorded, so the ratio is observable: CLOUD-475 records that a comment already counts as a durable home. Re-open predicate: re-open if the observation window shows spin-off comments rising while spin-off creates fall and fixes do not, which is the signature of the toll being routed around rather than paid or avoided. Acceptance
Not in this issue Deciding whether a given spin-off was legitimate — the judgement the gate must never make. The In Review transition gate, which is CLOUD-512's. And retrofitting receipts for branches predating the recorder, which is why the gate fails open on their absence. |
… was refined Every gate here prices FAILING to record something — finding-sink-check fails a turn that cites evidence and writes nothing, deferral-check fails a PR that defers without naming an issue. Nothing prices the opposite. Filing satisfies all of them in seconds while finishing costs a diff, a suite and a landing, so the punt is not a temptation, it is arithmetic. This is the sensor half and it gates nothing yet, which is normally the "log without a gate" non-negotiable 2 refuses. The exception is measured: the gate's firing rate cannot be estimated retrospectively, because which branch put which row on the board has never been recorded anywhere. The one available proxy — rows created between a PR opening and its merge — fires on 183 of 184 over the last 40 merged PRs, since a window over a fleet captures everyone's filings. So this record is what makes the gate specifiable, and its enabling trigger is a number this produces. The lint runs here rather than in a receipt the agent mints, and that is the whole design. ready-lint reads a payload the caller assembles: it was run three times during this issue's own refinement against text in a local file, once under the id CLOUD-NEW for a row that did not exist, green every time. A PostToolUse body has no such problem — it fires on the tool result, which is the tracker's response to the create. Three things were measured rather than assumed, and each would have shipped broken: The result envelope. No hook in this tree had ever read a tool result, and the documented example is a Write with a flat object. An MCP tool returns the content-block list, so `.tool_response.id` does not exist and a body written against the docs records nothing, silently. Relations. ready-lint's §8 rule cross-checks prose claiming a blocker against the payload's relations, and the create response carries none — so linting the response alone refuses exactly the rows refined most carefully. The create call's own blockedBy argument is the entire relation set on a create, so the payload is synthesised from both halves of the event. The body being judged is still the tracker's. How ready-lint is invoked. `mise run` needs the project cwd and a hook inherits the cwd of the tool call, so the obvious call fails outside the project — and mapping any non-zero to "unready" would record a refusal about the environment wearing the mask of a verdict about the row. Sibling task by path, and the status read as three answers: 0 ready, 1 unready, anything else unanswered. Pointer-only is load-bearing rather than decorative here: the text read is the entire issue body, and four fields reach the file. Refs: CLOUD-514
48d56f3 to
44fc7ec
Compare
|
|
/fast-forward |



Phase 1 of CLOUD-514: the sensor. It records what this branch put on the board and
whether a new row was refined when it was filed. It gates nothing yet.
Why a sensor alone, when non-negotiable 2 refuses a log without a gate
The exception is measured rather than asserted. The gate's firing rate cannot be
estimated retrospectively, because which branch put which row on the board has never
been recorded anywhere — not on the row, not in git, not on the PR. The one available
proxy (rows created between a PR opening and its merge) was measured over the last 40
merged PRs and fires on 183 of 184, since a window over a fleet captures everyone's
filings. So this record is what makes the gate specifiable at all, and phase 2's trigger
is a number this produces.
Why the lint runs in the hook rather than in a receipt
The first draft had the agent run
ready-lintand mint a receipt. That is worthless:ready-lintreads a payload the caller assembles, and during this issue's ownrefinement it was run three times against text in a local file — once under the id
CLOUD-NEW, for a row that did not exist — green every time. A toll payable in textnobody filed is not a toll. A
PostToolUsebody has no such problem: it fires on thetool result, which is the tracker's own response to the create.
Three things measured, each of which would otherwise have shipped broken
The result envelope. No hook in this tree had ever read a tool result, and the
documented example is a
Writewith a flat object. An MCP tool returns the content-blocklist, so
.tool_response.iddoes not exist — a body written against the docs recordsnothing, silently. The key is reached through
.text | fromjson | .id.Relations.
ready-lint's §8 rule cross-checks prose claiming a blocker against thepayload's relations, and the create response carries none. Linting the response alone
therefore refuses exactly the rows refined most carefully. The create call's own
blockedByargument is the entire relation set on a create, so the lint payload issynthesised from both halves of the event — body and
updatedAtfrom the tracker,relations from the input. The thing being judged is still the tracker's bytes.
How
ready-lintis invoked. Found by a failing test.mise runneeds the projectcwd and a hook inherits the cwd of the tool call, so the obvious call fails outside the
project — and mapping any non-zero to
unreadyrecords a verdict about the environmentwearing the mask of a verdict about the row. Sibling task by path, and the status read as
three answers:
0ready,1unready, anything else unanswered.Coverage
tests/board-write-record.batsat 14 rows, including the one this design turns on — arow whose §8 claims a blocker still records green, which is red without the relations
synthesis. Two
#MUTANTrows (create/update collapse, relations dropped);mutantgreenat 14/14 across ten gates. Full suite 1489/1489, run with
GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_SYSTEM=/dev/nullso the CLOUD-513 class ofverify/CI disagreement is caught here rather than on a runner.
Pointer-only is load-bearing rather than decorative: the text this reads is the entire
issue body, and four fields reach the file. Asserted.
Not in this PR
filed-here-checkand itslandcall site — phase 2, blocked on the observation windowabove.
Corrections, added after merge
Three claims above are wrong, and the evidence for each arrived within the hour.
"Unwired in the session that wrote it (CLOUD-187)." It fired immediately from the
working tree, recording five rows between 06:46 and 07:10 — before this PR merged at
07:08 — and two
PreToolUseguards another branch landed mid-session fired on this one.The measurement and its limits are on CLOUD-187.
Those first rows exposed a defect this PR shipped. A comment row recorded the
comment's own uuid rather than the issue key, because a
save_commentresponse is thecomment object and names no row at all. Sink 2's definition — a comment on the row that
already owns the finding — was therefore unobservable, while every count still looked
right. Fixed in the follow-up: the key comes from
.tool_input.issueId, and a reply or anon-issue parent records
-rather than a uuid, because a uuid in an issue-key column isa wrong answer wearing a right answer's shape.
The "1489/1489" figure was measured with
BATTEN_CLAIM_GUARD_BYPASS=1exported, sothat sweep ran against a tree with
claim-guarddisabled and five of its rows would havefailed.
test:batsinherits ambient bypass variables the same way it inherits ambient gitconfig; recorded on CLOUD-513, whose
[tasks."test:bats".env]block is where both belong.The number needs re-taking in a cleaned environment.
Refs: CLOUD-514