refactor(ci): retire derived-check, man-pages and ci-local-parity; repair the lap journal's basis - #770
Conversation
CLOUD-1145 `derived-check` is 289.8s — 23.8% of the corpus — and the comparator is already in the Rust tier for two of its four cells: retire it as SUBSUMED, carrying one set-equality assertion
Why
The cost is process-spawn tax, not compilation. RE-SCOPED (2026-08-30): the port was the wrong questionThis row was filed as a port — §1 said "the successor is a verb in A new verb widens a closed surface. House style §2's command list does not contain And §11 undercuts the predicate itself. Completions, man pages and markdown are derivations of that same runtime-emittable spec, so the shipped binary and the generated docs can never drift. This gate spends 210.8s — 16.9% of the whole bats corpus — checking for a drift the design says is impossible. Either the design holds and the comparator guards nothing, or it does not hold and the defect is that the artifacts are committed and comparable at all. So the question this row must answer FIRST, and it is one read: should If they must stay committed — a distribution argument, not a policy one — the home is an existing §2 verb: a The measurement below stands and is the reason this row is worth answering rather than deferring. What changes is that the port is the last option considered, not the first. The drag-in is its own header, not a callerNothing resolves But
Nothing outside the bundle resolves Three deleted paths, three ledger arms, zero governed sibling edits — and per the corrected audit the arms are **two ** Independent of
|
| key | why |
|---|---|
input.tree.documents |
the parsed workflows, mise.toml, release-plz.toml, renovate.json5 — properties 1–10, 13–17 |
input.tree.lines |
property 11's unquoted-# rule (:360-367) is only detectable pre-parse — the parse is what destroys the evidence. Also abandon-matrix.sh/land.sh read as text (:1081, :1088) |
input.tree.tracked / input.tree.missing |
property 12 (:695) asserts .github/dependabot.yml is absent, and the anti-vacuity guards at :649-659 need "did I see any workflow at all" |
Format::Yaml is one of the four parseable formats (crates/batten/src/facts.rs:1497-1508, :1538-1542 — Toml, Yaml, Json, Json5 parse; Pkl is declarable and never parsed), so the workflows resolve with no new parser.
One risk to validate BEFORE writing properties 8/10/15/16
**YAML 1.1 maps a bare **on: key to a boolean. Confirm how Format::Yaml renders the on: block into input.tree.documents before writing any predicate over a workflow trigger. The lines fallback is available if it bites. A predicate over a key the parse renamed is CLOUD-845's dead-gate class, and it would pass its own suite green.
Landability re-confirmed, in detail
**101 **@test declarations, 107 gate invocations. setup() (:10-67) spawns no cargo, no mise, no yq and builds no binary — every line is a builtin or a heredoc into $BATS_TEST_TMPDIR. So the cost is ~107 runs of a 1093-line bash program, each forking dozens of text utilities per fixture workflow. That is gate-side but behind no injectable seam — ci-local-parity.sh:259-263, :1030, :1031 are path overrides only — so none of the 54.6s is reachable from an ungoverned caller. It moves only if the suite dies.
Inbound references, ~90 tree-wide, all classified: exactly one task-name call (hk.pkl:608-610, ungoverned) and exactly one by-path resolution (tests/ci-local-parity.bats:14, which dies in this delta — and its $BATS_TEST_DIRNAME/.. spelling is CLOUD-1149's unrepointable shape, which does not matter because the file carrying it is deleted). Every other hit in a governed file is prose in a comment. Two governed files carry the bare token as fixture prose inside an issue-body payload (tests/ready-lint.bats:459, tests/spec-ref-check.bats:29) and need no edit at all.
$MUTANT_GATES: listed at mise.toml:457; four #MUTANT rows at ci-local-parity.sh:902, 903, 983, 984. All four re-home with the gate.
Verdict: RETIRABLE TODAY.
Why
Unit 13 of the 83-unit partition. 1 program, 1 suite, 46.3s — 3.7% of the 1244.6s serial suite.
CORRECTION (2026-08-29): this row is the SECOND-largest unblocked retirement, not the first
This row was filed claiming to be "the largest one that is blocked by nothing at all" and "the highest-value retirement available with no precondition". Both are wrong. CLOUD-1145 (derived-check + man-pages) is 210.8s / 16.9% — 4.5x this row — and carries no blockers on the board. Re-verified 2026-08-29: tests/derived-check.bats:2 subjects two governed programs, both deletable; there is no separate man-pages suite; and its only external read is cargo run -p batten at :106, which is the successor and therefore not a blocker.
The cause was a category error in the list below: it named derived-check as blocked because a row had been filed for it. A filed row is not a blocker, and CLOUD-1145's own §8 says "Blockers: None".
Dispatch order is CLOUD-1145 first, then this row. Together 257.1s / 20.7% with zero preconditions.
1093 lines, tests/ci-local-parity.bats 46.3s, in $MUTANT_GATES, 4 #MUTANT rows.
Both axes clear:
- Landable — no inbound reference from any governed file. Two deleted paths, two ledger arms, zero drag-in.
- Expressible — re-verified 2026-08-29 against the engine, not assumed.
Format::Yamlis parseable (crates/batten/src/facts.rs:1501,1540,1614), anddocumentstakes a glob (rules.rs:5368;declared_documentsreturns every match), so the successor sees a workflow no declaration names — which is exactly the case this gate exists to catch.tests/ci-local-parity.bats:2declares one governed subject, soSubjectFacts::diedis satisfiable by this delta alone.mise.toml:1857invokes it by task name;hk.pkl:610likewise. Both ungoverned.
That combination makes this row takeable with no precondition — second in line behind CLOUD-1145, which is also unblocked and larger. The units that ARE blocked: U1 needs CLOUD-1155's dissolution, hooks-wiring-check needs the out-of-root fact (CLOUD-1167, consumed by CLOUD-1160), U2 needs the same plus the forge producer (CLOUD-1159). derived-check is not among them.
HOME (2026-08-30): a PRESET and a consumer module, split by genericity
This row said the successor is *"a *policy/*.rego module" without saying whose. CLOUD-1176 makes that the first question, and this gate splits cleanly:
- GENERIC → a vendored PRESET. "No job runs on a draft", "every
pull_requestworkflow declaresconcurrencywithcancel-in-progress", "its trigger includesready_for_review" are CI hygiene any consumer using GitHub Actions wants. They name no Button identifier. Home: aci-hygienebundle besidetrunk-based,shell-hygiene,commit-hygieneandpinned-toolchainincrates/batten/src/policy/presets/**(CLOUD-836). - CONSUMER-SPECIFIC → this repo's
policy/*.rego. The$CI_REQUIRED_CHECKSroster cross-check and "every task CI runs is oneverifyruns" name this repository's job roster and task names. Non-negotiable rule 1 keeps them out of the core.
The split is not tidiness: a preset ships to every consumer, so putting the roster half there would bake Button's job names into everyone's binary — the exact violation §9's closing line names.
The one verdict to re-check before building
The census flagged ci-local-parity.sh:942 — mise tasks info test:cargo --json — as a possible "executes another program" blocker, then cleared it: the task block is in mise.toml, and the program's own comment at :956 says "The task's own line is in $manifest". So input.tree.documents answers it.
That clearance is the weakest link in this row and is stated rather than buried. If the successor needs mise's include and inheritance resolution rather than the literal block, the predicate flips to needing a resolved-task fact that does not exist. Settle that first — it is one read of what :942 actually consumes.
What the gate holds, so the port conserves rather than reinterprets
The properties that make CI a confirmation rather than a discovery: no job runs on a draft; every pull_request workflow supersedes its own runs; no job starts before the landing lease authorises its branch (CLOUD-420); and every task CI runs is one verify runs. Plus the $CI_REQUIRED_CHECKS roster cross-check — a pull_request job missing from the set, or a name matching no job, fails the gate (CLOUD-327).
zizmor.yml broke the first two for its whole life, so a draft touching a workflow still spent a runner and re-drafting did not close the tap. That is the defect class this gate exists for, and the port must keep catching it.
Refinement — Ready (clean on both axes; the mise-tasks-info precondition is settled)
Refinement gate: Definition of Ready & Done. This body carries only specializations.
- **Authority boundary (§1). **
mise-tasks/ci-local-parity.shandtests/ci-local-parity.batsare deleted — two paths, two// carried:arms. Two successors, per the home split above: aci-hygienePRESET bundle for the generic half and apolicy/*.regomodule in this repo for the roster half, each with acrates/batten/tests/*.rstier. **No new **crates/battenverb. The$MUTANT_GATESentry and 4#MUTANTrows move with it.mise.tomlandhk.pklcall it by task name and are ungoverned, so both repoint freely. - Computable predicate (§2), conserved not reinterpreted. Over the committed workflow set and
mise.toml: no job runs on a draft; everypull_requestworkflow declaresconcurrencywithcancel-in-progress; everypull_requestworkflow's trigger includesready_for_review(CLOUD-503 — a draft-eraopenedskip is a non-answer nothing can supersede); every task CI runs is oneverifyruns; and$CI_REQUIRED_CHECKSnames exactly thepull_requestjobs, in both directions. - The precondition, now SETTLED (§2). Whether
:942'smise tasks info test:cargo --jsonneeds only the literalmise.tomlblock or mise's resolved task graph. Answered in the block at the top of this row: the literal block suffices, and re-expressing it is a design decision that accepts a second authority over a body mise owns. Left unanswered it would have been CLOUD-845's dead-gate class. - Deliberately not in scope (§2). Changing
$CI_REQUIRED_CHECKSor the workflow set.CI_FANIN_CHECK's membership rule (CLOUD-900), which is conserved as-is. Deriving[ci].required_checks(CLOUD-54), which would change what the roster is — this port keeps reading the hand-maintained one. - Effect (§3).
read. It reads committed YAML and TOML and decides. - Output and exit (§5). Pointer-only: the workflow path, the job name, and which property failed — never a workflow body. Exit follows the
0/1/2/3table. - Commit / bump (§6).
refactor(ci)— no bump. Below0.1.0every release-worthy type collapses to a patch, butrefactoris not one: it releases nothing at any version. CLOUD-595's correction. - Test obligation (§7). Over the compiled binary in
crates/batten/tests/; **no **.batsfile is added or edited (V-SHELL-RULE-ADDEDrefuses one atdeny). **Two **// carried:arms, one per deleted path (CLOUD-908). Shown able to fail per CLOUD-418, four observed: a job runnable on a draft is reported; apull_requestworkflow withoutcancel-in-progressis reported; a$CI_REQUIRED_CHECKSname matching no job is reported, and apull_requestjob missing from the roster is reported — both directions, since CLOUD-327 is the one that reports green on a SHA nothing judged. Anti-vacuity: the current tree passes. Mutated: all 4 rows re-homed,mutant-censusgreen. Replayed: old and new over the same workflow set, agreeing finding-for-finding — CLOUD-1115 is the standing caveat on replay's tree arm. - Blockers (§8). None.
relatedToCLOUD-1151 (the wave owner), CLOUD-1140 (suite cost), CLOUD-1115 (the replay instrument), CLOUD-908, CLOUD-418.
Acceptance
- Both paths deleted with one ledger arm each; **no governed **
mise-tasks/*.sh**or **tests/*.batsedited. - The preset half names no Button identifier — asserted, since it ships to every consumer (rule 1, §9).
- No new
crates/battenverb;batten spec --format jsonemits exactly the committed row set. - The
mise tasks infoquestion is answered in writing before the module is built — done, see the block at the top of this row; what remains is that the successor states the second-authority trade it accepts by reading[tasks."test:cargo"]as a document, rather than silently re-deriving what:881-885warns about. - The
on:-as-boolean question is settled before any trigger predicate is written, and the answer is recorded — whichever way it falls. - All four properties still caught, both roster directions asserted.
mutant-censusgreen; 4 mutations honoured at their new home.bench/suites/RESULTS.mdregenerates with the suite absent, serial total down by the measured amount — ~54.6s on currentmain. Report it whichever way it falls.
Unit 13 of 83. The second-largest unblocked retirement — take it after CLOUD-1145, per the correction above. Seconds and blocker sets are two orthogonal axes; CLOUD-1174 owns that model and the artifact that makes this row's ranking checkable rather than asserted.
CLOUD-1218 The lap journal records a hand-emptied `target` as a WARM lap, so the ratchet inverts and the floor climbs to one nothing can satisfy
Why
CLOUD-1157's lap journal ratchets each floor to the worst consumption it has observed, which is right. What decides which floor a lap is charged to then mislabels a from-scratch rebuild as warm, so the warm floor learns a cold lap's demand.
CORRECTED 2026-08-30, third instance (PR #770). This paragraph originally read "what decides which floor a lap is charged to is whether the escalation dropped a basis-moving root." **That has not been true since **CLOUD-1157 (#756) landed —
basis_ofalready derives the basis from the TREE, not from the escalation:// crates/batten/src/prune.rs:1596 fn basis_of(root: &Path) -> Basis { let populated = directories_named(root, "deps") .iter() .any(|deps| std::fs::read_dir(deps).is_ok_and(|mut e| e.next().is_some())); if populated { Basis::Warm } else { Basis::Cold } }The residual defect is the
.any(), and the function's own doc comment already names it (prune.rs:1582-1589): *"this reads EVERY *deps*under the root, so a populated *target/release/depsreports warm while the DEBUG build the lap is about to run is cold." One populated profile masks another being empty.This matters because §2(b) below is written against the old mechanism and prescribes a fix the code already has. Reading it as written sends an author to replace an escalation check that is no longer there.
Measured on this container, 2026-08-30, while landing CLOUD-746. $GIT_DIR/batten-prune/laps.json held:
{"open":{"free_mb":9151,"basis":"warm","head":"d3788d66","measured":"2026-08-30"},
"ratchet":{"warm":{"mb":22861,"head":"5647a306","measured":"2026-08-30"},
"cold":{"mb":4640,"head":"090dd1f5","measured":"2026-08-30"}}}cold 4640 MB below warm 22861 MB is impossible — a cold lap consumes more than a warm one by definition, and [prune]'s declared numbers say so (warm 6242, cold 14914). The inversion is the tell, and it is the cheapest possible detector.
How it got there, and why an agent will keep doing it. A land lap that refuses on disk prints "Free space outside ./target, or start a fresh session." The obvious reading is rm -rf target, which is outside target-prune entirely. Each such lap then rebuilds from nothing and is charged to warm, ratcheting the warm floor toward the cold one. After a handful of laps the warm floor stood at 22861 MB — above anything a completed lap on this box can leave free — so every subsequent lap refused, in both directions: warm target failed the opening reading, cleared target failed the closing one.
The recovery is to delete the journal, which restores the declared floors. That is not discoverable from the refusal: the message names free space, the floor, and the roots it could not reclaim, and never that the floor it is quoting is one the tool taught itself.
SECOND INSTANCE (2026-08-30, landing PR #751) — and it defeats this row's own detector
Reproduced on a different container, different branch, same mechanism. Journal at refusal:
{"open":{"free_mb":8538,"basis":"warm","head":"6e0e7f6c","measured":"2026-08-30"},
"ratchet":{"warm":{"mb":15798,"head":"c96288c1","measured":"2026-08-30"},
"cold":{"mb":21519,"head":"c96288c1","measured":"2026-08-30"}}}Against declared warm 6242 / cold 14914, both floors are self-taught and both are ~2.5x and ~1.4x the declared value. Two hand-emptyings produced them, in this order:
| act | journal's reading |
|---|---|
rm -rf target/debug/incremental (the one root target-prune says it cannot reclaim) |
"consumed 15798MB — worse than any warm lap on record, so the observed warm floor rises to 15798MB" |
rm -rf target (after the above still refused) |
"consumed 21519MB — worse than any cold lap on record, so the observed cold floor rises to 21519MB" |
After that, every lap refused in both directions exactly as this row describes, and land failed three consecutive times.
Why §2(a) and acceptance bullet 1 would NOT have caught it
This ratchet is correctly ordered — cold 21519 > warm 15798 — so "a ratchet whose cold observation is below its warm one is refused" is silent here. The inversion is one symptom of poisoning, not the class. The recorded instance inverted because its two hand-emptyings happened to charge the wrong buckets; this one charged the right buckets and simply ratcheted both past what the box can satisfy. A detector keyed on ordering therefore catches the first instance and misses the second, which is the more ordinary shape — an agent following the refusal's own advice twice, in the order the refusal suggests it.
So the predicate wants a second conjunct, and §2(b) is the one that generalises: derive the basis from what the tree WAS at lap open. Under (b) both of my laps are cold and neither teaches the warm floor anything, which is the correct outcome and is reached without reference to ordering. Worth stating on the row because (a) is the cheaper check and reads as sufficient — it is not, and a fix that ships only (a) leaves this instance live.
A cheap third guard, orthogonal to both: refuse to ratchet a floor above the declared one by more than some factor without saying so. A learned number 2.5x its own config is more likely a mismeasurement than a real budget, and the refusal that quotes it should say which it is — which §5 already asks for and this instance makes concrete.
The recovery worked, and its cost is the real damage
rm -rf "$(git rev-parse --git-dir)/batten-prune" restored declared floors and admitted the next lap at 8130MB free against warm 6242MB. The cost was one full cold rebuild — the rm -rf target this row predicts an agent will reach for, which consumed 21519MB and ~20 minutes of wall clock, and which was never necessary: the journal was the whole problem and deleting it alone would have sufficed at the very first refusal. That is the measured price of the undiscoverable recovery, and it is the strongest argument for §5's "observed vs declared" wording landing with the fix rather than after it.
Why it is not CLOUD-861's or CLOUD-1030's. CLOUD-861 is the once-per-lap precondition, and the closing reading it added is what reports this — correctly, over a poisoned number. CLOUD-1030 is the escalation invalidating the basis that certified the lap. This is the third: an external reclaim mislabels the basis a lap is recorded under, so the ratchet learns the wrong thing and the error compounds across laps rather than affecting one.
Refinement — Ready
Refinement gate: Definition of Ready & Done. This body carries only specializations.
-
Authority boundary (§1).
crates/batten/src/prune.rs—LapJournal,OpenLap.basisandRatchet.[prune]'s declared floors are untouched: this row makes the recorded basis honest, it does not move a measured number. No new config key and no runner decides any part of it. -
Computable predicate (§2), REWRITTEN 2026-08-31 after both original clauses were implemented and refuted. One clause, decidable over the compiled binary: a profile cargo has BUILT whose
deps**is missing or empty makes the tree cold. **target/<profile>/.fingerprintis the marker — cargo writes one per profile it has built and leaves it behind whendepsgoes, so the absence becomes visible. Every.fingerprint's siblingdepsmust be populated; no.fingerprintanywhere falls back to the pre-existing reading, which is what keeps this a narrowing rather than a new requirement. It needs no guess about which profile the caller will build next, so the comment's refusal of profile scoping still stands.(a) A ratchet whose~~cold~~observation is below its~~warm~~one is refused.REFUTED. Silent on two of three instances (one journal carriedcold: null; one carried a correctly ORDERED pair). The strengthened form — refuse a warm observation at or above the declared cold floor — was implemented and turneda_warm_laps_consumption_does_not_raise_the_cold_floorred: that case deliberately drives a warm lap consuming 22000MB against a 14000MB cold declaration, because a warm observation under it cannot discriminate a per-basis ratchet from a shared one, and its comment records that as a surviving mutation. The premise is also a unit error — the declared floor is a free-space budget, not a ceiling on consumption — and no factor separates the real case (3.96x) from the fixture's legitimate one (3.67x).(b) …narrowing~~basis_of~~'s~~.any()~~to~~.all()~~.**REFUTED. **directories_namedonly yields directories that EXIST, and the reclaim an agent actually performs REMOVEStarget/debug/deps, which drops that profile out of the walk and leaves a populatedtarget/release/depssatisfying either quantifier..all()only helps for adepsthat survives but is empty, which is not what any of the three instances did."Empty or near-empty
root" was also wrong and would miss all three. Mine lefttarget/release,target/debug/buildand ~660 MB standing; the second instance deleted onlytarget/debug/incremental.depsis the build basis; the rest of the tree is not. -
Effect (§3). Unchanged.
target-prunekeeps its classification; what changes is which bucket an observation lands in. -
Output & exit (§5). Unchanged, except that the refusal should be able to say the floor it quotes is observed-and-inverted rather than declared — the current message gives a reader no way to tell a learned floor from a configured one. Pointer-only throughout: megabytes, a head and a date, as today.
-
Test obligation (§7). Over the compiled binary, shown able to fail per CLOUD-418. The discriminating case needs TWO profiles:
a_tree_emptied_by_something_other_than_the_escalation_is_still_a_cold_onedeletesdepson a one-profile tree, so the walk comes back empty and any quantifier answers cold — it cannot see this class. Red before: a built profile whosedepsis REMOVED while another profile stays intact. Green and staying green: every built profile intact is still warm, and a tree cargo never fingerprinted still takes the fallback — without that pair the fix is satisfied by charging everything tocold, which raises the COLD floor instead and fails the same way one bucket over. -
Blockers (§8). None.
relatedToCLOUD-1157 (whose journal this is), CLOUD-861 (the closing reading that surfaces it), CLOUD-1030 (the escalation-invalidates-basis half) and CLOUD-1153 (the other way this refusal misreports its own cause).
Acceptance
- ✅ A profile cargo has built whose
depsis missing or empty makes the tree cold, whichever other profiles survive. Landed on refactor(ci): retire derived-check, man-pages and ci-local-parity; repair the lap journal's basis #770 (856c3746). - ✅ A tree whose every built profile is intact is still warm, and a tree cargo never fingerprinted is judged by the pre-existing reading. Landed on refactor(ci): retire derived-check, man-pages and ci-local-parity; repair the lap journal's basis #770.
- ✅ The refusal names the journal that holds a learned floor, so the recovery is discoverable from the message rather than costing a full cold rebuild to find. Landed on refactor(ci): retire derived-check, man-pages and ci-local-parity; repair the lap journal's basis #770 (
1c896f89). - ~~A ratchet with ~~
~~cold~~~~below ~~~~warm~~~~never decides a refusal. ~~Withdrawn — see §2. Silent on two of three instances, and the strengthened form conflicts with a landed measured test. - ~~A correctly-ordered but inflated ratchet is also caught. ~~Withdrawn with it: no threshold separates the real case from the fixture's legitimate one. With the basis now honest, a cold lap is charged to the cold ratchet and the warm one is never taught a rebuild's demand, which is the route §2 takes instead.
- Still open, and the reason this row is not Done: the ratchet remains unbounded, so a genuinely mismeasured observation of any basis is still learned permanently. That needs a predicate over satisfiability — a floor above what a completed lap can leave free can only ever refuse — rather than over ordering or magnitude. Not attempted here.
Found while landing CLOUD-746: six consecutive land laps refused on a floor the tool had taught itself from my own rm -rf target, and the fix was deleting a file no message named.
CLOUD-1216 A path glob cannot select a step for a file the commit DELETED, so `suite-bench-check` — whose whole predicate is set equality over `tests/*.bats` — is silent on exactly the change that breaks it
Why
suite-bench-check decides one thing: that bench/suites/RESULTS.md's membership equals git ls-files 'tests/*.bats', in both directions (mise-tasks/suite-bench-check.sh:68-82). The reverse direction — a corpus row naming a suite the tree no longer carries — is refused as "a cost attached to nothing".
Its hk step is globbed on bench/suites/RESULTS.md and tests/*.bats (hk.pkl), and hk.pkl's own comment states the intent exactly:
Globbed on the corpus and on the suites, because both directions rot: a suite added without regenerating is a file nothing records, and a suite deleted leaves a cost attached to nothing.
The glob cannot deliver the second half. hk selects steps by matching changed paths against the glob, and a deleted path is not there to match. So the one commit shape the reverse direction exists for — a deletion — is the one shape that does not select the step.
Measured, 2026-08-30, on PR #770
Commit f45e214 deleted tests/derived-check.bats as part of CLOUD-1145's retirement. It passed the full pre-commit gate. bench/suites/RESULTS.md still recorded that suite, so the tree was left in the exact state suite-bench-check exists to refuse:
::error:: suite-bench-check: bench/suites/RESULTS.md records tests/derived-check.bats,
which is not a tracked suite — a cost attached to nothing.
That state survived four further commits, every one of which also passed the gate, and was found only by running the task by hand. Nothing in the gate would have caught it before CI.
The severity is that the failure is silent and delayed rather than that the corpus was stale. The corpus going stale is cheap and obvious once seen. What this row is about is that a gate can be correct, registered, globbed with deliberate intent stated in a comment, and still structurally unable to fire on half the class it names.
The class is wider than this one step
Any step whose predicate is about a file's absence, or about set equality against a tracked-file listing, inherits this. A glob over paths answers "did one of these change", and a deletion is a change the glob cannot see. Steps whose predicate is purely about the content of surviving files are unaffected, which is why this has not surfaced before.
Related but distinct: CLOUD-899 is a glob naming a path that never existed (wrong glob), and CLOUD-949 wants hk's effective plan as a pre-admission fact (visibility). This is neither — the glob is right, the plan is right, and the selection still cannot happen.
Refinement — Ready (make the gate's selection independent of the deleted path)
Refinement gate: Definition of Ready & Done. This body carries only specializations.
-
**Authority boundary (§1). **
mise-tasks/suite-bench-check.sh's own selection contract, and the hk step that invokes it. The predicate itself is correct and is not in scope — this changes only whether it runs. No other step is modified in this row; the wider class is surveyed in the acceptance and filed separately if it has members. -
Computable predicate (§2). The step must run on any commit that changes the set of tracked
tests/*.bats, including one that only removes members. Two candidate mechanisms, and the row does not pre-judge between them: either the step declares no path glob and always runs (it is milliseconds —hk.pklalready records that it "re-runs NOTHING"), or hk's selection is fed the deleted paths as well as the surviving ones. The first is decidable today and needs nothing from upstream; the second is the general fix and may not be expressible inhk.pkl. -
Deliberately not in scope (§2). The set-equality predicate, its exit codes, and the corpus format. Timing is deliberately ungated and stays so.
-
**Effect (§3). **
read. The step reads two committed listings and compares them. -
Output and exit (§5). Unchanged: pointer-only, naming the suite and the corpus, over the existing
0/1/2contract. -
**Commit / bump (§6). **
ci— no bump; it changes which commits a gate's step is selected for, and the crate releases nothing at any version. Notfix:ready-lintstrips the scope before reading the type (sed -E 's/[(][^)]*[)]//'), sofix(ci)declaresfix, which implies a patch and contradicts the "no bump" on the same line. -
Test obligation (§7). Shown able to fail per CLOUD-418, and the discriminating case is the one this row exists for: a commit whose ONLY change is deleting a
tests/*.batsmust select the step and turn it red while the corpus still records that suite. Green and staying green: a commit touching neither the corpus nor any suite must not pay for the step if the always-run route is taken — measure it, sincehk.pkl's "re-runs nothing" claim is the whole affordability argument.**
⚠️ **hk check --planCANNOT ANSWER THIS AND WILL TELL YOU THE ROW IS WRONG. Measured 2026-08-31: it reportssuite-bench-checkeven for a commit touching onlycrates/**/*.rs, which neither glob matches — so it does not respect globs, and old-config and new-config runs come back identical. Three selection probes against it "refuted" this row and I began correcting the body before the control caught it.The probe that decides is a commit. Stage a deletion of a tracked
tests/*.batswhile the corpus still records it, and try to commit: with the glob the commit SUCCEEDS and the step never runs; without it the commit is REFUSED, naming the stale row. Run both ways ontests/land.bats. -
Blockers (§8). None.
Acceptance
-
✅ A commit that deletes a tracked
tests/*.batswithout regenerating the corpus is refused at the pre-commit gate, not at CI. Landed on refactor(ci): retire derived-check, man-pages and ci-local-parity; repair the lap journal's basis #770 — the step is unglobbed, and the deletion-only probe flips from committing cleanly to being refused. -
✅ The step's cost on an unrelated commit is measured and stated: 531ms, three runs (531/530/531). Stated in
hk.pklbeside the step, replacing a comment that said "milliseconds" — which a reader takes as ~10. -
✅ The wider class is surveyed, and
suite-bench-checkis its only member, so no further rows are owed. Done on refactor(ci): retire derived-check, man-pages and ci-local-parity; repair the lap journal's basis #770.Method, so a reader can judge it rather than take it. The class is not "reads a tracked listing" — 19
mise-tasks/*.shreadgit ls-filesand almost all of them use it to ENUMERATE subjects, where deleting a file simply removes a subject and creates no violation. The class is narrower: a predicate where a path's ABSENCE is itself the refusal. Two passes:- The 19
git ls-filesreaders, checked for directionality.module-map-checkis representative of the majority — it refuses "modules absent from the map" only, so a deleted module cannot make it red (its map row goes stale un-gated, which is a different and lesser thing). - A search across
mise-tasks/for the reverse-direction refusal shape — "which is not tracked", "no longer exists", "attached to nothing", "names no file". One task refuses on it:suite-bench-check. The other three hits (reference-check,closing-key-check,hooks-wiring-check) carry the phrase in comments about unrelated staleness, not in a refusal over a tracked listing.
The bound, stated rather than implied: this surveys
mise-tasks/. A step whose predicate lives in a Rego module or the crate is not covered, and a member phrasing its refusal differently would be missed. Both are cheap to re-run if a second instance ever appears. - The 19
Found while retiring derived-check (CLOUD-1145) — the deletion passed the gate and the corpus stayed stale across four commits.
CLOUD-1233 The escalation drops `incremental` and the very next build regenerates it, so a container near its allowance thrashes — measured 6859MB per lap against 613MB with `CARGO_INCREMENTAL=0`
Why
[prune] declares incremental as a regrowable, basis-moving root, and the escalation drops it when a lap opens below the warm floor. That buys space exactly once: the next cargo build writes it straight back, so the lap ends below the floor again and the following lap escalates again. On a container whose allowance is close to the working-set size the cycle does not converge.
Measured, 2026-08-31, one container, same branch and same HEAD
Consecutive mise run land laps over 3005244d, differing only in CARGO_INCREMENTAL:
| lap | escalation dropped | lap consumed | free at close | warm floor |
|---|---|---|---|---|
| default | 5393MB | 6859MB | 7037MB | 7264MB |
| default | 581MB | 746MB | 7023MB | 7264MB |
CARGO_INCREMENTAL=0 |
— | 0MB | 7071MB | 7264MB |
CARGO_INCREMENTAL=0 |
239MB | 613MB | 7050MB | 7264MB |
11x on lap consumption, and the 5393MB the escalation reclaimed in the first lap is almost exactly what the build then rewrote. The reclaim and the regeneration are the same bytes going round.
Note what the refusal says while this happens: "escalated below the warm floor — 5393MB of regrowable cache dropped; none of those roots is the cargo build's basis, so the next build is still warm and the warm floor is what applies." That is true about the basis and misleading about the budget — the drop does not reduce what the next lap will consume, it guarantees the next lap re-spends it.
Why this is not CLOUD-861's or CLOUD-1030's
CLOUD-861 is the once-per-lap precondition and the closing reading that reports the overspend — it is the sensor that makes this visible, correctly. CLOUD-1030 is the escalation invalidating the basis that certified the lap, which is about which floor applies. This is about the consumption term: the escalation's own reclaim is the largest single input to the next lap's cost, and nothing accounts for that.
It is also distinct from CLOUD-1157's coverage work. Widening the reclaim to more classes does not help here; incremental is already reclaimed, and reclaiming it is what creates the cost.
The shape of a fix, not pre-judged
Two candidates, and the row does not choose between them:
- Do not write what you are about to drop. If a lap is admitted near the floor, run its builds with incremental compilation off. The gate already knows it is near the floor — that is the escalation's own trigger — so the information is in hand at the right moment.
CARGO_INCREMENTAL=0is the measured lever above and costs some rebuild time in exchange for not thrashing. - Stop treating
incrementalas free headroom. If dropping it reliably costs ~5GB on the next lap, it is not regrowable-at-no-cost in the sense the escalation assumes; the accounting should say so, and the escalation should prefer roots whose regeneration is not the next lap's largest term.
The first is decidable today and needs nothing from upstream. The second is the honest model and may want [prune] to carry a regeneration cost per root beside the existing cold flag.
Refinement — Ready (stop the reclaim from funding the next lap's overspend)
Refinement gate: Definition of Ready & Done. This body carries only specializations.
- **Authority boundary (§1). **
crates/batten/src/prune.rs's escalation, and whichever surface sets the build environment for a lap.[prune]'s declared floors do not move. No change to which roots are declared regrowable. - Computable predicate (§2). A lap that escalates does not consume, on its next build, materially what the escalation just reclaimed. Decidable as a comparison of two recorded numbers the journal already has —
escalated_mbon one lap againstconsumedon the next — so the acceptance is measurable rather than argued. - Deliberately not in scope (§2). The floors themselves, the basis decision, and the ratchet. Those are CLOUD-1030's, CLOUD-1218's and CLOUD-1197's respectively.
- **Effect (§3). **
readplus the existing reclaim. Setting a build environment variable for a lap is not a new authority; it is the same process the lap already spawns. - Output & exit (§5). Pointer-only, unchanged. If the escalation stops claiming that a dropped root is free, its line should say what it actually costs — the current wording implies the opposite.
- Test obligation (§7). Shown able to fail per CLOUD-418, over the compiled binary. Red before: a lap that escalates, followed by a lap whose consumption is ~the escalated amount. Green and staying green: a lap that does not escalate is unaffected and still writes its incremental cache — without that twin the fix is satisfied by disabling incremental compilation everywhere, which is a real cost paid on every container including the ones with room.
- Blockers (§8). None.
Acceptance
- The escalation's reclaim is not the dominant term in the next lap's consumption.
- A lap with headroom is unaffected and keeps its incremental cache.
- The escalation's report does not describe a root as free when regenerating it is the next lap's largest cost.
- The measured pair above is the regression case: same HEAD, same branch, 6859MB against 613MB.
Found while landing PR #770: four consecutive laps refused at the closing reading, each having just reclaimed roughly what it then re-spent.
CLOUD-934 The inline-regex refusal exempts presets, so the anti-duplication mechanism does not apply to the code that ships to every consumer
Why
CLOUD-885 made a pattern's home decidable: an inline regex in a policy module is refused at load, so a pattern must be a [[pattern]] row read as data.batten.patterns["<id>"]. pattern.rs states the value plainly — duplication becomes unwritable rather than merely detectable (measured: one concept, 19 spellings across 17 shell programs), and the pattern inventory becomes reviewable data (house style §11).
check_no_inline_regex returns early when rule.preset.is_some() (policy.rs:1269). So the refusal does not apply to presets — and presets are the code that reaches every consumer, while a consumer module reaches one repository.
Measured: crates/batten/src/policy/presets/shell-hygiene/sibling-resolves.rego writes six inline regexes (:39, :53, :71, :77, :95, :100), legally, including a name_capture whose trailing character class carries a documented false-positive fix. Nothing holds those six to the inventory [[pattern]] exists to be.
Why the exemption exists, and why that reason does not settle it
The exemption is not arbitrary: a [[pattern]] row is consumer config, and a vendored preset ships with no consumer config to reference. A preset naming data.batten.patterns["x"] would be a module with a dangling reference in every repository that had not declared x — which is worse than an inline regex, because it is a dead gate rather than an unindexed pattern.
So the exemption is correct as far as it goes. What is missing is the third option: presets have no inventory of their own. The question this row exists to settle is whether a preset bundle can carry its own pattern table — vendored alongside the modules, compiled in by the same include_str! table (policy.rs:229-261), invisible to consumer config and therefore not a dangling reference — or whether inline is genuinely the right answer for a preset and the exemption should say so where a reader will find it.
Why a row and not a comment on CLOUD-885
CLOUD-885 is In Review. A comment on a row about to land is a finding with no reader, and the exemption is a distinct decision from the one that row made — 885 decided where a consumer's pattern lives.
Refinement — Ready
Refinement gate: Definition of Ready & Done. This body carries only specializations.
- Source of truth (§1). Whatever the verdict, there is one inventory per bundle scope: consumer patterns in
batten.toml, and either a vendored preset table or an explicit recorded decision that presets are exempt and why. Not both, and not silence. - Computable predicate (§2). Conditional on the verdict, and both branches are decidable: if presets get a table,
check_no_inline_regexstops exempting them and a preset carrying an inline regex is refused at load. If inline stays correct for presets, the predicate is that the exemption is documented at its site — whichrules-driftcannot hold, so that branch ships a test assertingpolicy.rs:1269's early return carries a reason, in the shapespawn_census.rsuses againstclippy.toml. - Effect (§3).
read. No new authority either way; a vendored table isinclude_str!at build time, which is the existing preset mechanism rather than a new one. - Generated artifacts (§4).
schema/batten.schema.jsononly if a consumer-visible key appears — a vendored table should not add one, and that is part of what makes it the safer branch.derived-checkgates it. - Output & exit (§5). Unchanged. A load-time refusal is exit 1 — a config fault, not a policy verdict (house-style §7) — which is what
check_no_inline_regexalready does for consumer modules. - Commit / bump (§6).
fix(policy)— patch until0.1.0. Tightening a refusal to cover presets could in principle break a downstream preset author, but there are none: presets are vendored and compiled in, so the blast radius is this repository's own four modules. - Test obligation (§7). Shown able to fail per CLOUD-418. If presets get a table: a preset with an inline regex is refused at load, and
sibling-resolves.rego's six become table references with the module's owntest_rules still green — the migration must be behaviour-preserving, asserted by those tests rather than claimed. If inline stays: the reason atpolicy.rs:1269is asserted present, and the test is shown able to fail by deleting it. - Blockers (§8). None.
relatedToCLOUD-885 (which built the refusal and the[[pattern]]table), CLOUD-836 (which vendored the presets this exempts).
Acceptance
- A verdict is recorded: presets get their own vendored pattern inventory, or inline is correct for a preset and the reason is at the exemption's own site.
- Whichever branch, a mechanism holds it — a load-time refusal or an asserted rationale, never prose alone.
sibling-resolves.rego's six inline patterns are either migrated with itstest_rules still green, or explicitly covered by the recorded verdict.
ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Free Run ID: ⛔ Files ignored due to path filters (1)
📒 Files selected for processing (21)
💤 Files with no reviewable changes (5)
🚧 Files skipped from review as they are similar to previous changes (12)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughAdded the Merge Risk: ⚪ Minimal · up to The PR retires legacy CI gates and moves their checks into maintained test and policy paths while repairing journal basis handling; no actionable merge-blocking risk remains beyond normal checks and review. Note 🎁 Summarized by CodeRabbit FreeYour organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/settings/billing. Comment |
|
Standing down on the SonarQube failure: it is not this PR's, and nothing here reads it.
So this is the documented steady state of the repository rather than a regression introduced here, and there is no fix to port: the row that would change it is open and blocked on a separate remediation. Two corrections to my own earlier reasoning, recorded because both were wrong in the direction of ignoring a signal:
No action taken on it. Everything in the required roster is untouched by this observation. Generated by Claude Code |
86556a4 to
d6dad71
Compare
1c896f8 to
fa49c3b
Compare
485da84 to
2b4be4e
Compare
`tests/derived-check.bats` was 289.8s — 23.8% of a 1219.4s bats corpus and its
single most expensive suite — spent re-answering a question
`crates/batten/tests/surface.rs` already answers over the compiled binary. The
disposition is SUBSUMED rather than a port: `:66` and `:222` held the drift half
for both artifact families before this change, so almost none of the 289.8s was
coverage.
What was genuinely missing is the SET half, and it was missing in OPPOSITE
directions per family, because each side's expected set was anchored on a
different authority:
* completions — expected set is the fixed `SHELLS` const, so a
declared-but-uncommitted script failed at `fs::read` and an EXTRA committed
file was never looked at.
* man — expected set is `committed_pages()`, which reads the directory, so an
extra file had to render and match and a page the surface DECLARED with no
committed file was invisible.
`the_committed_artifacts_are_exactly_the_ones_the_surface_declares` closes all
four cells with one assertion rather than patching two, which is what
`derived-check.sh`'s `comm -23` reverse scan did. It compares PATHS and never
opens a file, so non-negotiable rule 4 holds structurally. The man half derives
from `batten spec --format json`; the completions half stays anchored on
`SHELLS` because the spec carries the `--shell` flag but not its value set —
stated in the code rather than faked as a derivation.
`committed_pages()`'s doc comment was false in two directions and is rewritten
rather than repointed. It called `mise-tasks/man-pages.sh` "the one authority
for which pages exist": that script was itself a derivation of `batten spec`, so
it was never an authority, and the function never read it, so the crate
described a derivation that did not happen. Retiring the script is what
surfaced it.
`[tasks.man]` now derives its page list inline from the same spec — the `jq`
program `man-pages.sh` already was — so a hop that read as a second authority
and cost a `cargo run` per invocation is gone.
MEASURED
* the new assertion passes on this tree in 0.01s against the suite's 289.8s
* shown able to fail in all four cells (CLOUD-418), each naming the artifact:
man declared-not-committed, man committed-not-declared, completions
declared-not-committed, completions committed-not-declared. The last two are
the cells the existing tier was blind to and the ones this row exists for.
* `mise run man` regenerates all 77 pages to a zero-byte diff
* `mise run mutant-census` green: 104 gates, all declared or exempt
No `#MUTANT-EXEMPT` row is owed. `mutant-census.sh` collects subjects from
`mise-tasks/*.sh` and `policy/*.rego`, both one level deep; this retirement
creates no module, so there is no source to be `uncovered`. Dropping
`derived-check` from `$MUTANT_GATES` is the whole obligation, without which the
census reports `names-no-subject`.
No governed file was edited: three deleted, and every surviving reference is in
ungoverned config or prose.
Refs: CLOUD-1145
…l-parity
First of two commits retiring `ci-local-parity` (54.6s, 4.5% of the bats
corpus). This one lands the successor for the half that is true of the PRACTICE
rather than of this repository; the consumer-specific half and the deletion
follow.
FOUR PROPERTIES, and the split from the consumer's own module is non-negotiable
rule 1 rather than tidiness. "No job runs on a draft", "a pull-request workflow
supersedes its own runs", "every workflow declares a concurrency group" and "a
draft-gated workflow subscribes to ready_for_review" name no repository, no
task and no job. A required-check roster and a bot's branch prefix do, and a
preset reaches every consumer — shipping them here would bake this repository's
job names into everyone's binary.
ONE PROPERTY IS RE-SCOPED, deliberately and in the safe direction. The retired
program scoped `ready_for_review` to "produces a required check", which is a
consumer's fact. Draft-gating is the same condition read from the workflow
itself: a job that skips on a draft is one whose verdict can only arrive on the
ready event. The preset asks it of workflows that draft-gate, so it stays
computable without a roster.
TWO PARSE QUESTIONS THE SHELL DID NOT HAVE, both settled over the compiled
binary rather than by reading:
* `on:` — YAML 1.1 resolves a bare `on` to a boolean, which would key the
trigger block as `true` and make every trigger predicate a dead gate,
passing its own suite green while deciding nothing. This engine's parser
resolves only `true`/`false`, so `on` arrives as written.
`the_engine_keys_the_trigger_block_as_on` proves it by DISCRIMINATION — the
fixture can only be refused if the block was read as `on` — rather than by
inspecting a fabricated document, which is the one thing a `with input as`
case cannot do.
* `cancel-in-progress` — the shell matched it as TEXT, where the quoted and
bare spellings are identical. Through a parser they are a string and a
boolean, and only the boolean is the flag GitHub honours. A module written
against `"true"` would be green in its own suite and dead in the field.
MEASURED
* `mise run policy-test`: 24 bundles, 274 passed, 0 failed.
* `crates/batten/tests/ci_hygiene.rs`: 12 passed in 0.71s, including
`this_repository_is_clean_today` over all 24 committed workflows — the
anti-vacuity term, and what says a preset that refused everything would not
look identical to one that discriminates.
* Shown able to fail (CLOUD-418): a job runnable on a draft, a workflow that
never supersedes, a workflow with no concurrency group, and a draft-gated
workflow that can never be superseded. Each names the workflow.
* The shapes that must NOT be refused are asserted too: a scheduled workflow
declining to cancel its own next tick, and a workflow that does not
draft-gate not being asked for the ready subscription.
A defect found in this file's own first draft, recorded because the failure mode
is silent: `object.union` is a DEEP merge, so two removal fixtures built by
overriding a parent key kept the key they meant to remove and were byte-identical
to the clean one. Both deny cases passed nothing. They are spelled out as whole
documents now.
`schema/batten.schema.json` and its override twin regenerate with the new preset
name: the enum is derived from `policy::preset_names()`, so the committed schema
is what refuses a preset that does not exist.
The preset ships no `#MUTANT-EXEMPT` row and needs none: `mutant-census`
collects `mise-tasks/*.sh` and `policy/*.rego`, both one level deep, and a
vendored preset lives under `crates/batten/src/policy/presets/`.
Refs: CLOUD-1161
Second of three commits retiring `ci-local-parity`. The generic half shipped as
the `ci-hygiene` preset; this lands the half that names things only this
repository has, so the deletion can follow with a successor already live for
every property.
WHAT IS HERE, and why it could not go in the preset. A required-check roster, a
task called `verify`, a fan-in called `final`, the ecosystems this tree
maintains, the two shell programs that make the abandon safe — every one is a
consumer fact. A preset reaches every consumer, so shipping them there would
bake this repository's job names into everyone's binary, which is the violation
non-negotiable rule 1 names.
Five predicates, thirteen refusal classes: every task CI runs is one `verify`
runs (with the foreign-runner exemption, which is an allowlist of FOREIGN labels
rather than of Linux ones — exempting anything not `ubuntu-*` would switch the
property off for `self-hosted` and for a matrix expression); the roster names
exactly the pull-request jobs, in both directions; the release PR opens as a
draft; the retired second bot stays retired and the surviving one keeps the five
bounds that decide what its lane spends and covers, with its commit type written
where a preset cannot outrank it; and the fan-in is named, homed, read rather
than restated, and actually called.
THE `verify` HOP IS SPELLED OUT rather than chased. `verify` became a
dependency-free exit-code mapper under CLOUD-407 and the gate set moved to
`verify:gated`, so a reader following only the first hop reports every one of
those tasks as CI-only — a false alarm rather than a missed one, but one that
fires on every commit. Following `mise run` calls transitively would be a second
authority on the task graph mise owns, which is the same objection that keeps
the cargo-spelling predicate out of this file entirely.
MEASURED
* `mise run policy-test`: 25 bundles, 295 passed, 0 failed.
* `crates/batten/tests/ci_parity.rs`: 17 passed in 3.55s, including
`this_repository_is_clean_today` — the case that says every predicate here
holds against the real roster, release config, Renovate config, dependabot
absence and fan-in wiring, rather than only against fixtures.
* Shown able to fail (CLOUD-418) for all thirteen classes, and the two the row
names specifically are both directions of the roster: a job missing from it
and a name matching no job.
* The shapes that must NOT be refused are asserted too: a foreign runner
running a task `verify` does not, and a matrix leg matching on its base name.
* `mise run mutant-census` green: 105 gates.
TWO REFUSALS THE ENGINE MADE ME TAKE, both correct and both recorded because
they are the load-time tier doing its job. A Rego function cannot be
multi-valued, so reading a job's tasks as `ci_task(path, name)` faulted at
evaluation and is a partial set instead. And an inline regex is refused at load,
so `mise run <task>` is a `[[pattern]]` row — which is the registry preventing
exactly the duplication the retired program had, where the same concept was
spelled once in an `awk` pipeline and again in its own prose.
One case of my own first draft asserted the wrong property: it put a task name
inside a `run:` body and expected it to be ignored, which the retired program
would also have matched. Parsing subsumes the harder half of that property —
a YAML comment does not survive into the document at all — so the case now pins
what is actually decided, that the reading is bounded to `run:` scalars.
Refs: CLOUD-1161
…perties
The other half of what a hosted-CI run's configuration has to satisfy. Where
`spend-is-authorised` asks what a run COSTS, these ask whether the wiring that
decides it does anything at all: a trigger no job admits, a filter written where
filtering is already too late, two schedules on one minute, a fan-in asserting
three of its four dependencies, a cache-warm compile guarded on a step that does
not exist. Every one is silent when it breaks — the run list looks normal and
the conclusion is green.
Seven predicates, eight refusal classes, all generic: they name no repository,
no directory and no task, so they belong beside the four already in this bundle
rather than in the consumer's module.
TWO ARE SCOPED NARROWLY ON PURPOSE, because the wide reading fires on correct
configurations:
* A declared trigger is only asked to be admitted where a job condition
MENTIONS the event name at all. A workflow that does not discriminate by
event answers for every trigger it declares, so judging it would refuse
ordinary workflows.
* A fan-in is only asked to assert its whole dependency set where it already
names SOME of them. A job that names none is not making a claim about its
`needs:` — it simply waits, which is what `needs:` is for.
MEASURED
* `mise run policy-test`: 25 bundles, 315 passed, 0 failed.
* `crates/batten/tests/ci_hygiene.rs`: 20 passed in 3.12s, including
`this_repository_is_clean_today` — so all eleven properties in this bundle
hold against the real 24 committed workflows, not only against fixtures.
* Shown able to fail over the compiled binary for the shapes a fabricated
document cannot prove: a cron SEQUENCE spanning two files, a `needs` ARRAY,
the trigger map iterated by key, and a step id nested in a sequence read by
a string in a sibling mapping.
* The discriminating halves are asserted too: a staggered cron pair, a
guarded cache-warm compile, and a `workflow_run` with no branch condition at
all — the last because a deliberately repository-wide trigger is not a
defect.
A DEAD GATE THE SECOND TIER CAUGHT, and the load-time tier structurally could
not. `ci_hygiene.rs` passed `patterns: &[]`, so the rule reading
`data.batten.patterns["cache-hit-step-id"]` had an undefined pattern, never
fired, and BOTH its deny case and its clean case passed. That is CLOUD-845's
class arriving inside the test harness rather than in the module: a
`[[pattern]]` row is the consumer's data, so unlike a preset's verdicts it is
not supplied by the binary, and a fixture that omits it silently disables every
rule that reads one. The tier supplies it now, and the deny case fails without
it.
`cache-hit-step-id` is a `[[pattern]]` row rather than an inline literal even
though presets are currently exempt from that refusal — `.claude/rules/policy-modules.md`
calls the exemption a hole rather than a design, so this bundle does not take it.
Refs: CLOUD-1161
Three more of `ci-local-parity`'s properties, all consumer-specific because each
names something only this repository has: a step called "Landing lease
precondition", a task called `checks-green`, and the branch prefixes two bots
land on.
* NO JOB SPENDS BEFORE THE LEASE AUTHORISES ITS BRANCH. The lease serialises
landing, but enforcing it only inside the lander means anything else pushing
to a ready pull request buys a full matrix without ever touching the lock —
measured as four concurrent matrices while the lease changed hands three
times, every holder honouring it. The step must be FIRST: a job that
installs a toolchain and then asks permission has already spent most of what
asking was meant to save. Jobs that wait on others are exempt for a REASON
rather than by enumeration — they cannot start ahead of the cancellation.
* AND THE PRECONDITION MUST BE ALLOWED TO FAIL. The presence clause matches
the step's NAME, so a copy that reds its own job reads as present and
correct. Counted rather than searched for absence, because the two forms
differ only by the suffix and a bare search passes a file carrying both.
* A WORKFLOW READING CHECK STATUS DECIDES GREEN THROUGH ONE PREDICATE. Every
hand-rolled copy so far has counted a wholly skipped set as zero
outstanding, i.e. green — and a wholly skipped set is exactly what a
draft-era refresh looks like. Keyed to the ENDPOINT rather than to a banned
spelling, since the spelling is what a rewrite changes.
* EVERY LIVE BOT LANE HAS A WATCHER AT ITS TRIGGER. Nothing runs on a bot's
behalf unless a workflow watches its heads, so handing a lane to a bot
without a lander is a complete, silent failure. The prefix is read from the
config that OWNS each lane rather than assumed, and a lane whose config is
absent is not asked for a watcher.
A FOURTH DEAD GATE THE SECOND TIER CAUGHT, and the third of this kind in this
PR. `V-LEASE-PRECONDITION-FATAL` reads `input.tree.lines`, and the row's
`line_sources` declared only the two shell programs — so no workflow was ever in
that fact and the rule could not fire at all. Its load-time case passed, because
a `with input as` case fabricates the `lines` object the row would never build.
The glob is declared now and the case fails without it.
That is three separate instances in this change of the same class: a predicate
whose input the row does not actually supply. Each was invisible to the
load-time tier by construction and each was caught by driving the compiled
binary, which is exactly the argument `.claude/rules/policy-modules.md` makes
for the second tier existing.
MEASURED
* `mise run policy-test`: 25 bundles, 325 passed, 0 failed.
* `crates/batten/tests/ci_parity.rs`: 22 passed in 7.16s, including
`this_repository_is_clean_today`.
* Shown able to fail over the compiled binary for all four: a job with no
lease step, a lease step that is not first, a precondition invoked without
the tolerant suffix, a workflow rolling its own green predicate, a live lane
with no watcher, and a prefix named only in a job condition — the last
because a condition is evaluated after the run exists, so it is not a scope.
Refs: CLOUD-1161
Deletes `mise-tasks/ci-local-parity.sh` (1093 lines) and
`tests/ci-local-parity.bats` (1674 lines, 101 cases, 54.6s — 4.5% of the bats
corpus). Every one of its 40 predicates now has a live successor, landed in the
preceding commits, so nothing is ungated across the deletion.
THREE HOMES, split by what each predicate can honestly name:
* `ci-hygiene` PRESET — what a run costs and whether its wiring can be reached.
Generic: no repository, no task, no job name.
* `policy/ci-parity.rego` — the roster, the task graph, the release and bot
configs, the fan-in. Consumer facts, kept out of the core by rule 1.
* `[tasks."cargo-spelling"]` — the foreign-runner second spelling, which needs
mise's own answer about mise's task graph. A module cannot spawn, so the only
expression available there was a second parser over a body mise already owns
— the objection the retired program stated in its own words. It stays a mise
task, so that objection is honoured rather than paid.
This commit also lands the last generic predicate: an unquoted `#` that swallows
a `${{ }}` interpolation. Read from `lines` and not the document, which is forced
rather than chosen — the parse is what DESTROYS the evidence, so by the time the
value is a node it is already truncated and nothing downstream can tell a short
string from a swallowed one. Anchored rather than counted, because the obvious
raw-vs-parsed count was measured at 75% false positives.
THE LEDGER — 101 case arms and 2 path arms, and the dispositions are not
uniform, which is the whole point (CLOUD-908):
* carried 86 — the assertion moved home.
* subsumed 6 — every one is the PARSE doing the work. A comment does not
survive into a parsed document and a parsed value is one
shape whatever its source formatting, so the classes the
shell excluded by hand cannot arise here at all.
* changed 6 — two are the `ready_for_review` re-scope (a roster is a
consumer fact and cannot live in a vendored preset, so the
preset scopes on draft-gating instead, which is the same
condition read from the workflow itself). Four are
could-not-look cases: the engine separates absent from
unparseable through `input.tree.missing`, where the shell had
one channel for both and had to refuse the empty case to
avoid a vacuous pass.
* withdrawn 3 — the success-line cases. All three assert the retired
program's own stdout summary. A policy module emits findings
and says nothing on success (house style §6), so there is no
summary line for a successor to carry. The visibility they
bought is now the rules' `governed` guards plus
`this_repository_is_clean_today`, which fail loudly instead
of reporting a count.
Claiming `carried` over a predicate nobody ported would have been exactly
CLOUD-908's recorded failure, so each of the twelve non-carried arms names what
diverged and why rather than being rounded up.
`$MUTANT_GATES` drops the entry; the four `#MUTANT` rows re-home into
`policy/ci-parity.rego` alongside its `#MUTANT-EXEMPT CLOUD-1161`, which is owed
because `mutant` resolves a gate's suite as `tests/$gate.bats` and
`V-SHELL-RULE-ADDED` refuses creating one. `hk.pkl`'s step is replaced by
`cargo-spelling`'s and three stale prose references are repointed.
`cargo-spelling` is shown able to fail in all three arms, both could-not-look
cases included: a drifted task spelling, a task yielding no cargo invocation, and
a tree with no foreign-runner subject at all. A gate that found nothing must not
look like a gate that passed.
Not in this commit: `bench/suites/RESULTS.md`. It is regenerated last, after
`test:bats` runs over the deleted tree — `suite-bench.sh` records the hazard of
doing it the other way round, where a report older than the tree still names a
retired suite.
Refs: CLOUD-1161
…ubject
A regression my own preset introduced, caught by the full bats run rather than
by any tier I wrote: `tests/prebuilt-lint.bats`'s case
"a prebuilt install-action step is not a violation" went red on
`.github/workflows/t.yml workflow-declares-a-concurrency-group`.
WHY IT REACHED THAT SUITE AT ALL, which is the part worth recording. The retired
gate was a TASK pointed at one directory, so it only ever judged the real tree.
Ported into a config rule, the predicate now travels with `batten.toml` into
every fixture tree that copies it — and `prebuilt-lint`'s deliberately does,
because its header says the fixture must judge "this commit's engine and this
commit's config as the pair that ships". That widening is inherent to the port
and I did not anticipate it.
That fixture is therefore a canary for a property worth having: THE SHIPPED
RULESET MUST NOT REFUSE AN ORDINARY MINIMAL REPOSITORY. An unconditional "every
workflow declares a concurrency group" does refuse one, and for a preset that
ships to every consumer that is too much.
THE NARROWING, and it is principled rather than a fudge. A group is required
where two runs in flight are two ANSWERS TO ONE QUESTION — `pull_request`,
`issue_comment`, `workflow_run`, `schedule` — the same pull request, the same
comment thread, the same upstream run, the same recurring job. A `push`-only
workflow is the one case where that does not hold: each run is keyed to a
DIFFERENT commit, so two runs are two subjects rather than two answers, and
superseding is a cost preference rather than a correctness property.
Every measured instance of the original defect is inside the narrowed set. The
one the property was built for — N concurrent comment invocations running N
concurrent attempts to advance the trunk, at 245 refusals against 6 merges in
half an hour — is `issue_comment`. What is given up is push-only coverage, and
this tree has no push-only workflow lacking a group, so no live verdict moves.
The ledger arm moves with it: "the concurrency property judges every workflow,
not only the pull_request ones" is now `// changed:` rather than `// carried:`,
naming what narrowed and why. Rounding it up to carried would have been the
CLOUD-908 failure this PR exists to avoid — 85 carried, 7 changed, 6 subsumed,
3 withdrawn.
MEASURED
* `mise run policy-test`: 25 bundles, 330 passed, 0 failed.
* `crates/batten/tests/ci_hygiene.rs`: 20 passed, including
`this_repository_is_clean_today` — the real 24 workflows still satisfy the
narrowed rule, so nothing in this tree changed verdict.
* `tests/prebuilt-lint.bats` "a prebuilt install-action step is not a
violation": green.
* A push-only workflow is now asserted clean, which is the discriminating case
the narrowing exists for.
The other failure in that run, `config-deprecations`'s "no release tag to
compare against; fetch tags", is this container having zero tags and is not
this change.
Refs: CLOUD-1161
`suite-bench-check` is bidirectional set equality against `git ls-files 'tests/*.bats'`, so the corpus had to be regenerated from a report that post-dates the two deletions. Regenerated rather than hand-edited, per the file's own instruction: `mise run test:bats` first, then `mise run suite-bench --write`. 142 suites, 797.3s serial on this machine. MEASURED, as a RATIO INSIDE ONE RUN — which is the only comparison the data supports. `main`'s own corpus, one report on one machine, records: tests/derived-check.bats 221.8s 15.4% tests/ci-local-parity.bats 53.6s 3.7% ------------------------------------------ removed 275.4s 19.1% of that run's 1440.9s So the retirement removes ~19% of the serial corpus. The cost is removed rather than moved: the successors are `crates/batten/tests/surface.rs`'s set-equality assertion (0.01s), two policy modules whose tiers run under `test:cargo`/`policy-test`, and one `mise.toml` task. TWO CORRECTIONS TO THE FIGURE THIS COMMIT FIRST CARRIED, both of which made the result look better than the evidence allows. First, it subtracted across machines: 1219.4s (the issue body's measurement) minus 866.2s (this container) was reported as a -353.2s saving, and the 8.8s by which that overshot the billed 344.4s was explained away as run-to-run noise. The two numbers never shared a baseline, so neither the difference nor the explanation meant anything. `main`'s corpus recording 1440.9s for a LARGER suite count is the proof. The same objection applies to 1440.9s vs 797.3s here, which is why that subtraction is not stated either. Second, the share is smaller than CLOUD-1145/CLOUD-1161 predicted — 19.1% against the 28.2% those rows derived from a 1219.4s, 145-suite corpus. Nothing regressed: the corpus grew to 1440.9s across 144 suites between that measurement and now, so the same two suites are a smaller fraction of a bigger whole. Stated because a retirement whose headline number shrinks is exactly the result worth not omitting. Refs: CLOUD-1145
…the registry
Two of the preset's predicates were DEAD. `cache-warm-compile-is-guarded`'s
missing-step-id arm and `interpolation-is-not-swallowed` both read their regex
from `data.batten.patterns[...]`, which a preset can never resolve.
`policy.rs`'s own exemption comment says why, and says it is not a gap to be
closed:
"A preset is compiled in; a consumer cannot add a `[[pattern]]` row on its
behalf, and the preset cannot read one — so refusing it would make a
vendored bundle unloadable with no fix available."
So the lookup was undefined for every consumer, Rego read undefined as "does
not hold", and both rules gated nothing while reporting clean. A dead gate and
a clean tree are byte-identical on the decision surface.
THE HARNESS WAS HIDING IT, which is the part worth recording. An earlier
revision of `ci_hygiene.rs` declared those three ids in the fixture's own
`Vocabulary` — so the compiled-binary tier, whose whole purpose is proving the
engine builds the input the predicate reads, was supplying input no consumer
supplies. Its deny cases passed for the wrong reason. That table is now `&[]`,
which is what a consumer hands a preset, and it is what fails these cases if
anyone respells a literal as a registry lookup again.
SHOWN ABLE TO FAIL, both arms:
registry lookup + no consumer patterns -> `every_shipped_preset_passes_its_
own_suite` fails, naming test_a_guard_naming_a_missing_step_id_is_refused
and test_an_unquoted_hash_that_swallows_an_interpolation_is_refused
inline literal + no consumer patterns -> preset suite and the compiled tier
both pass
Rule 1 still binds the three literals and they hold: two are YAML's own comment
and interpolation syntax, one is GitHub Actions' `steps.<id>.outputs.cache-hit`.
None names a consumer.
The three `[[pattern]]` rows are dropped from `batten.toml` as orphans.
`mise-run-task` stays — `policy/ci-parity.rego` is an in-repo consumer module
and the registry is correct for it.
Refs: CLOUD-1161
`inline-task-bodies-not-growing` counts `run = '''` in mise.toml against origin/main. This branch raises it by one, 31 -> 32, and the body is `cargo-spelling` — the single predicate of `ci-local-parity`'s forty that did not become Rego. NEITHER ROUTE THE ROW'S OWN `no_fix_reason` NAMES IS OPEN AS A FIX HERE. "Migrate the predicate onto a rule kind" is refused by `RuleKind::scopes`, which pairs every spawning kind with `RuleScope::Tree` alone: `cargo-spelling` reads `test:cargo`'s effective body through `mise tasks info`, so a module would have to parse mise.toml a second time — a second authority over a task graph mise already owns, which is the retired program's own stated objection and not something to override while porting it. The other home, a `mise-tasks/*.sh` file task, does not avoid a ratchet either: it grows `bash-surface-not-growing` instead, and adds net-new authored shell to the exact surface CLOUD-843's campaign — and this branch — is retiring. That is the same increase one row over, bought by running the campaign backwards. So waiving is the row's other sanctioned answer, taken deliberately rather than as an escape from it. EXPECTED TO LAPSE UNUSED, as the `tests-not-deleted` waiver says of itself: `base = "origin/main"`, so the floor becomes 32 the moment this lands and this row suppresses nothing thereafter. It exists to get one commit past the gate, not to stand. The two-week expiry is the mechanism — CLOUD-1137 is the row that would give this predicate a counted home, and a lapse makes someone re-read the decision instead of inheriting it. Refs: CLOUD-1161
…cy engine The last of `ci-local-parity`'s forty predicates, and the one CLOUD-1161 planned to keep as a mise task. Two gates closed that route, and neither is worked around here: `inline-task-bodies-not-growing` counts `run = '''` in mise.toml against origin/main, and a new task raised it 31 -> 32. `config-lint` then refused the waiver — correctly. Its admission is a `Weakens` clause groomed into the issue BEFORE the claim, copied into the branch's claim receipt at claim time; nothing an author writes at PR time reaches it. The claim predates the clause, and re-claiming to mint a fresh receipt is refused too, since `claim-check` exits non-zero on `assigned` and `has-pr`. So the task is gone and the predicate is `foreign-cargo-is-the-declared- spelling` in `policy/ci-parity.rego`. `run = '''` is back to 31, which is what removes both findings rather than suppressing either. WHAT THE PORT COSTS, stated rather than dissolved. The retired program objected that a second reader of the task body is a second AUTHORITY over a graph mise owns, and that objection is real. It is affordable here only because the two readings are the same bytes: `test:cargo` carries no template, no `depends` body and no argument substitution, so `mise tasks info --json`'s `.run` IS the manifest string. The bound is exactly that, and `V-TASK-CARGO-UNREADABLE` is the arm that surfaces the day it stops holding. A module cannot spawn (`RuleKind::scopes` pairs every spawning kind with `RuleScope::Tree`), so the alternative was not a better reading — it was no gate at all. Faithful to the shell in the three ways that decide cases: line-based across workflow files rather than job-based, `--no-run` exempt from the comparison AND outside the anti-vacuity term, and could-not-look refused on both sides rather than reported clean. BOTH FIXTURES GAINED THE SUBJECT, because neither carried one. The module's `sound_input` had no workflow lines and the compiled fixture had neither a `test:cargo` task nor a foreign leg, so the new rule would have been inert over every existing case — clean because nothing was looked at. That vacuity is only visible in the compiled tier: `line_sources` failing to declare the workflows leaves every deny case passing green. THE COMPILED TIER NEEDED A VERDICT-BEARING HELPER. `findings` returns `(path, line)`; the token is not on `Finding` at all, because `Violation` carries it and by the time a `Finding` exists it is gone (CLOUD-1120). Every case here is a different token over the SAME file, and the `--no-run` case has to show one firing while another does not, so the file's existing "is anything refused" shape cannot discriminate them. `verdicts_raised` reads `Scan::classes` instead. Seven `conserves` case arms named `mise.toml` as their successor and now name `crates/batten/tests/ci_parity.rs`. Left alone they would have been `carried` arms pointing at a task that no longer exists, which is CLOUD-908's recorded failure. The module doc block asserted "it stays a mise task" in two places and is corrected. policy test: 25 bundles, 334 passed, 0 failed. cargo: 3225 passed, 0 failed. Refs: CLOUD-1161
…learned CLOUD-1218's last acceptance bullet, and the only part of that row this change is honest enough to close. A refusal already says a floor is "observed on <head> rather than declared", which tells a reader the number was learned and still leaves them nowhere to go: it lives in a file no message mentions. The omission has a measured price. On another container an agent read the refusal, reached for the `rm -rf target` it DOES name, and paid a full cold rebuild — 21519MB and ~20 minutes — when deleting the journal alone would have sufficed at the first refusal. On this one I did the same thing one directory down and wedged the loop for hours. Pointer, not payload (rule 4): the path, only when the floor in force is observed rather than declared. TWO LARGER REPAIRS WERE ATTEMPTED AND BOTH ARE REFUTED. They are recorded in `basis_of`'s doc comment and on the row rather than left for the next author to rediscover, because each looks obviously right until it is run. The first was to bound the ratchet: refuse a WARM observation at or above the declared COLD floor, on the reasoning that a warm lap cannot cost more than a full rebuild. `a_warm_laps_consumption_does_not_raise_the_cold_floor` refutes it — that case deliberately drives a warm lap consuming 22000MB against a 14000MB cold declaration, because a warm observation UNDER the cold declaration cannot discriminate a shared ratchet from a per-basis one, and its comment records that as a surviving mutation. The premise also confuses two units: the declared floor is a free-space budget, not a ceiling on consumption. The second was to narrow `basis_of` from `.any()` to `.all()`, so one emptied profile makes the tree cold. That does not reach the measured instance: `directories_named` only yields directories that EXIST, and the reclaim an agent actually performs REMOVES `target/debug/deps`, which drops that profile out of the list and leaves a populated `target/release/deps` reading warm under either quantifier. A test written for it went red and is what proved this. What would close it is a signal for a profile that SHOULD carry `deps` and does not — a claim about cargo's layout this function does not currently make, and not one to guess at inside a row that is already someone else's. Refs: CLOUD-1218
…hers survive CLOUD-1218, and the repair the row exists for rather than the reporting half. Three containers in two days wedged their own landing loop here: a lap refuses on disk, the agent reclaims inside `target`, the rebuild that follows is recorded as a WARM lap, and the warm ratchet learns a full rebuild's demand — after which every lap refuses against a floor nothing on the box can satisfy. WHY THE OBVIOUS NARROWING DOES NOT REACH IT, stated because I shipped it first and it went red. `basis_of` read the `deps` directories that EXIST and asked whether ANY was populated. Requiring EVERY one changes nothing for the measured case: the reclaim an agent actually performs REMOVES `target/debug/deps`, so that profile drops out of the walk entirely and the surviving `target/release/deps` satisfies either quantifier. A removed directory is invisible to a question asked over the directories that are there. `.fingerprint` IS WHAT MAKES THE ABSENCE VISIBLE. Cargo writes one per profile it has built and leaves it behind when `deps` goes, so "is there a profile that has been built and now has nothing to build on" is answerable from the tree. Every `.fingerprint`'s sibling `deps` must be populated, or the next build for that profile writes everything — which is what `Basis::Cold` means. It needs no guess about which profile the caller will build next, which is the guess this function refuses to make and still refuses. And it is the same KIND of claim about cargo's layout that looking for `deps` at all already is — one directory over — not a new authority. That is where I first talked myself out of it, wrongly. NO `.fingerprint` ANYWHERE FALLS BACK to the older reading, so a tree cargo has never fingerprinted is judged exactly as before. That keeps this a narrowing rather than a new requirement: every fixture that writes a bare `deps`, and any consumer whose layout this does not describe, is untouched. The direction is what makes it safe: it can turn a warm reading cold and never the reverse, and cold is the stricter floor, so the failure mode is a lap held to a larger budget than it needs rather than one taught a number nothing can satisfy. The first is a delay; the second is the wedge. Shown able to fail (CLOUD-418), with the discriminating case the existing coverage cannot express: `a_tree_emptied_by_something_other_than_the_ escalation_is_still_a_cold_one` deletes `deps` on a ONE-profile tree, so the walk comes back empty and any quantifier answers cold. The new case keeps a second profile intact, which is what made the old reading say warm. Both anti-vacuity twins are there — every built profile intact is still warm, and a never-fingerprinted tree still takes the fallback — because without them the repair is satisfied by calling every tree cold, which raises the COLD floor instead and fails the same way one bucket over. NOT DONE HERE, and recorded rather than left: the ratchet still learns whatever a lap reports. Bounding it was tried and refused — `a_warm_laps_consumption_ does_not_raise_the_cold_floor` deliberately drives a warm lap above its cold declaration, because that is the only shape distinguishing a per-basis ratchet from a shared one, and the declared floor is a free-space budget rather than a ceiling on consumption. With the basis now honest, a cold lap is charged to the cold ratchet and the warm one is never taught a rebuild's demand, which is the route the row's own §2(b) says generalises. Refs: CLOUD-1218
…eachable there `.claude/rules/policy-modules.md` told every module author that the preset exemption from the inline-regex refusal is "a hole rather than a design" and to "write the row" anyway. Following that produces a DEAD GATE, which is strictly worse than the duplication the registry exists to stop. A `[[pattern]]` row is consumer config. A preset is compiled in and reaches a consumer who wrote no rows, so `data.batten.patterns["x"]` resolves to undefined there, Rego reads undefined as "does not hold", and the rule decides nothing while loading clean. `policy.rs` says so at the exemption's own site — the demand is unsatisfiable, because a consumer cannot add the row on a preset's behalf and the preset cannot read one. CLOUD-934 predicted this in those words. CLOUD-1161's `ci-hygiene` preset on this same PR is it happening: three registry reads, two predicates dead for every consumer, `policy test` green at 330 passed over them. Only `crates/batten/tests/policy_presets.rs` caught it, because it runs a preset's suite the way a consumer gets it. So the paragraph now says what the engine does, and carries the two things a reader needs beside it: that rule 1 still binds the literal (which is what makes inline safe rather than merely necessary — a preset ships everywhere, so its pattern could not name a consumer anyway), and that a compiled tier must hand the preset an EMPTY vocabulary or its own deny cases pass for the wrong reason. My harness declared the ids and hid the defect from the tier whose whole purpose is finding it. Whether presets should get a vendored inventory of their own stays CLOUD-934's open question; this only stops the file prescribing the one shape that cannot work today. Prose, no mechanism: `rules-drift` holds the input-key lists in this file to the generated schemas and reaches nothing here. Stated rather than implied, per the file's own §"What this file does not gate". Refs: CLOUD-934
… deletion CLOUD-1216. `suite-bench-check` decides set equality over `git ls-files 'tests/*.bats'` in BOTH directions, and its hk step was globbed on the corpus and on the suites for exactly that reason — the comment said so. The glob could only ever deliver the first half: hk selects a step by matching CHANGED PATHS, and a deleted path is not there to match, so the one commit shape the reverse direction exists for was the one shape that did not select the step. Measured on this PR: a commit deleting `tests/derived-check.bats` passed the full pre-commit gate and left the corpus recording a suite the tree no longer tracked. It survived four more commits, all green, and was found by running the task by hand. SHOWN ABLE TO FAIL THROUGH THE REAL PATH, because the obvious probe does not discriminate. `hk check --plan` reports this step even for a commit touching only `crates/**/*.rs`, which neither glob matches — so the plan does not respect globs, and three selection tests run against it "refuted" the premise and were themselves worthless. I nearly corrected the row on that basis. The probe that decides it is a commit: stage a deletion of a tracked `tests/*.bats` while the corpus still records it, and try to commit. with the glob the commit SUCCEEDS, the step never runs without the glob the commit is REFUSED, naming the stale row Run both ways on `tests/land.bats` here. THE COST, measured rather than asserted, because the comment this replaces said "milliseconds" and a reader reads that as ~10: 531ms, three runs (531/530/531). That is paid on every commit now, and hk runs steps concurrently so the marginal wall-clock cost is usually less. It buys a direction that was structurally unreachable. The wider class — any predicate about a file's ABSENCE, or about set equality against a tracked listing — stays CLOUD-1216's to survey; this closes the one step that demonstrably lost a verdict. Closes CLOUD-1216
CLOUD-1218's third failure mode, and it is what blocked this branch from landing across five consecutive `land` laps. `lap()` computed one basis and used it for two different questions: which ratchet bucket owns the lap's consumption, and which floor its remaining free space must clear. The first is rightly the basis the lap opened under — what a lap cost is a fact about the build that ran. The second is not: the floor asks whether there is room for the build that comes NEXT. Measured. A lap opened on an empty `target/` — correctly Cold — and closed on a fully built tree with 14839MB free. It was refused against the 17357MB cold floor while `verify` had SUCCEEDED inside it: a full cargo build, the whole cargo suite, all 2491 bats cases, no disk error anywhere. That floor is unreachable by construction at the close, because building the tree is precisely what spends the headroom a cold floor demands, so a cold-started lap refuses at its own close forever however much is reclaimed first. Every `rm -rf target` recovery opens exactly that lap. The tree's reading alone cannot decide it, and the first attempt at this repair got that wrong. `basis_of` reads `deps`, so a lap opened Cold because an earlier escalation dropped `incremental` has a full `deps` and reads Warm while its next build really is a full one — the OR that function's own header describes, whose second half the tree is blind to. `a_warm_laps_consumption_does_not_raise_the_cold_floor` caught it. So the discriminator is what the tree read AT OPEN, which `OpenLap` now carries beside the effective basis. A lap that opened on a cold tree and ends on a warm one has built it; one that opened on a warm tree under a standing escalation has changed nothing the tree can show. The field is optional, so a journal written before it parses rather than resetting a clone's lap history, and its absence falls back to the previous reading. CLOUD-861's opposite spiral is preserved by the same rule rather than in spite of it: a close whose own reclaim escalated opened warm, so it is judged warm. Refs: CLOUD-1218
A defect in this branch's own `.fingerprint` reading, found when `verify` refused its own precondition against a full rebuild's floor on a tree that had just built cleanly. `basis_of` walks the configured root unbounded, so it reaches every build tree NESTED inside it. This repository's suite writes its fixtures under `target/tmp/<case>/`, and one of them exists precisely to model a profile whose `deps` was removed. That fixture sits at `target/tmp/<case>/target/debug/.fingerprint` with an empty sibling, so "every built profile must have something to build on" found it and judged the whole repository cold. The predecessor had the same exposure and hid it: an `.any()` over fixture `deps` directories that happen to be populated read the tree warm, so the litter masked rather than refused. Neither is a reading of this build. Cargo writes `<root>/<profile>/` and `<root>/<triple>/<profile>/`, so the walk's results are filtered to those two depths. Anything deeper is a different tree that happens to live here. Refs: CLOUD-1218
`prune` reached 101 lines against the workspace's 100-line ceiling. The escalation is the cohesive unit to lift out: it is entered on one condition, owns its own two-tier rationale, and advances exactly the two readings the caller's accounting is built on. No behaviour change — `tests/target_prune.rs` is green across all 61 cases either side of the move. Refs: CLOUD-1218
2b4be4e to
2bdeb78
Compare
|
❌ The last analysis has failed. |
|
/fast-forward |
The rebase onto current `main` conflicted here because PR #770 landed and regenerated the same file for its own three retirements. A generated artifact is not hand-merged, so the rebase took `main`'s copy and this regenerates it with `mise run suite-bench --write` over a real `test:bats` run: 2443/2443 cases, 138 suites. Header moves to 138 suites / 1097.1s, and none of this bundle's four suites appears. That total is a whole-corpus re-measurement on this machine and is NOT a savings figure — the claim this bundle makes is the 65.7s its four suites cost against the 1440.9s / 144-suite corpus they were recorded in. Refs: CLOUD-1164
`target-prune` refused every `land` lap on this branch, and not for space: 21923MB free against a 17357MB cold floor, while `[prune.warm]` reported "measured against a tree that no longer exists — declared 128, live 140, tolerance 10". No receipt written, so no lap could complete. This branch caused it. It adds three stems under `crates/batten/tests/*.rs` (`reference_coverage.rs`, `skill_contract.rs`, `config_deprecations.rs`) and PR #770 added more on `main`, taking the live count 12 past a tolerance of 10. That is the gate working: the floor it defended was taken against a smaller tree, and a floor taken against fewer stems fails silently — the check passes, the build writes more than the basis anticipated, and exhaustion arrives as a rustc IO error inside a test run. `batten.toml`'s own block records the identical precedent on 2026-08-30, where a nine-suite bundle re-measured for the same reason. `count` and `measured` move together, as that block requires, and both floors scale by the per-stem model it already states: warm 56.7 x 140 = 7938, cold 135.6 x 140 = 18984. `mb` is validated to equal `worst_mb * multiplier` at load, so both fields move on each. 140 rather than 141: `git ls-files 'crates/batten/tests/*.rs'` crosses `/` and catches `tests/common/mod.rs`; the gate's selector does not, and 140 is the reading the comparison actually uses. NEITHER FLOOR IS AN INDEPENDENT MEASUREMENT, which the block warns matters. An honest cold number needs a build from an empty `target` and an honest warm one a minimal post-prune tree; these are the declared-value-scaled-by-stems derivation the block sanctions, and a reader needing either exact should take it rather than trust the scaling. Refs: CLOUD-1164
The rebase onto current `main` conflicted here because PR #770 landed and regenerated the same file for its own three retirements. A generated artifact is not hand-merged, so the rebase took `main`'s copy and this regenerates it with `mise run suite-bench --write` over a real `test:bats` run: 2443/2443 cases, 138 suites. Header moves to 138 suites / 1097.1s, and none of this bundle's four suites appears. That total is a whole-corpus re-measurement on this machine and is NOT a savings figure — the claim this bundle makes is the 65.7s its four suites cost against the 1440.9s / 144-suite corpus they were recorded in. Refs: CLOUD-1164
`target-prune` refused every `land` lap on this branch, and not for space: 21923MB free against a 17357MB cold floor, while `[prune.warm]` reported "measured against a tree that no longer exists — declared 128, live 140, tolerance 10". No receipt written, so no lap could complete. This branch caused it. It adds three stems under `crates/batten/tests/*.rs` (`reference_coverage.rs`, `skill_contract.rs`, `config_deprecations.rs`) and PR #770 added more on `main`, taking the live count 12 past a tolerance of 10. That is the gate working: the floor it defended was taken against a smaller tree, and a floor taken against fewer stems fails silently — the check passes, the build writes more than the basis anticipated, and exhaustion arrives as a rustc IO error inside a test run. `batten.toml`'s own block records the identical precedent on 2026-08-30, where a nine-suite bundle re-measured for the same reason. `count` and `measured` move together, as that block requires, and both floors scale by the per-stem model it already states: warm 56.7 x 140 = 7938, cold 135.6 x 140 = 18984. `mb` is validated to equal `worst_mb * multiplier` at load, so both fields move on each. 140 rather than 141: `git ls-files 'crates/batten/tests/*.rs'` crosses `/` and catches `tests/common/mod.rs`; the gate's selector does not, and 140 is the reading the comparison actually uses. NEITHER FLOOR IS AN INDEPENDENT MEASUREMENT, which the block warns matters. An honest cold number needs a build from an empty `target` and an honest warm one a minimal post-prune tree; these are the declared-value-scaled-by-stems derivation the block sanctions, and a reader needing either exact should take it rather than trust the scaling. Refs: CLOUD-1164
Retires two governed shell gates by porting their predicates and then deleting the shell, repairs the disk-floor defect that wedged this branch's own landing loop, and closes a gate that could not fire on the change it exists to catch.
Closes CLOUD-1145
Closes CLOUD-1161
Closes CLOUD-1218
Closes CLOUD-1216
Closes CLOUD-1233
DO-NOT-CLOSE CLOUD-934
~19% off the serial bats corpus. Both programs and both suites are gone; every predicate they held has a live successor with a compiled-binary test behind it.
The measured delta, stated as a ratio inside one run
An earlier revision of this body claimed 344.4s / 28.2%. That was a cross-machine subtraction and is withdrawn: the 1219.4s baseline came from the issue bodies and the post-retirement total from this container, and the two never shared a machine.
main's own corpus recording 1440.9s for a larger suite count is the proof.The comparison the data supports is a ratio within a single report —
main's corpus, one run, one machine:tests/derived-check.batstests/ci-local-parity.batsLower than the 28.2% CLOUD-1145/1161 predicted, and nothing regressed: the corpus grew to 1440.9s across 144 suites between those rows being written and now, so the same two suites are a smaller slice of a bigger whole. Stated because a retirement whose headline number shrinks is the result worth not omitting.
CLOUD-1145 —
derived-check+man-pagesDisposition SUBSUMED, not a port:
crates/batten/tests/surface.rsalready held the drift half over the compiled binary. What was missing was the SET half, and it was missing in opposite directions per family, because each side's expected set was anchored on a different authority:completions/man/SHELLS)the_committed_artifacts_are_exactly_the_ones_the_surface_declarescloses all four cells with one assertion. It compares paths and never opens a file, so non-negotiable rule 4 holds structurally.A false claim the retirement surfaced:
committed_pages()'s doc comment calledman-pages.sh"the one authority for which pages exist". Wrong twice — the script was itself a derivation ofbatten spec, and the function never read it. Rewritten to say what it does.Measured: the new assertion passes in 0.01s · shown able to fail in all four cells ·
mise run manregenerates all 77 pages to a zero-byte diff.CLOUD-1161 —
ci-local-parityThe issue body describes 5 properties; the program implements 40, across 101 test cases. All 40 are ported, into two homes:
ci-hygienepreset (11) — what a run costs and whether its wiring can be reached. Names no repository, task or job.policy/ci-parity.rego(29) — roster, task graph, release and bot configs, lease, fan-in, and the foreign-runner cargo spelling. Consumer facts, kept out of the core by rule 1.cargo-spellingis in Rego, not a mise task — the plan changed under two gatesThe 40th predicate was to stay a
[tasks.…]block, honouring the retired program's own objection that a second reader of the task body is a second authority over a graph mise owns. Two gates closed that route:inline-task-bodies-not-growingcountsrun = '''inmise.tomlagainstorigin/main; a new task raised it 31 → 32.config-lintthen refused the waiver, correctly. Its admission is aWeakensclause groomed into the issue before the claim and copied into the branch's claim receipt at claim time — nothing an author writes at PR time reaches it. The claim predates the clause, and re-claiming is refused too (claim-checkexits non-zero onassignedandhas-pr).So the task is gone and the predicate is
foreign-cargo-is-the-declared-spelling.run = '''is back to 31, which removes both findings rather than suppressing either. The cost is stated rather than dissolved: it readstest:cargo's body from the manifest where the shell readmise tasks info. Those are the same bytes only while that task carries no template, andV-TASK-CARGO-UNREADABLEis the arm that surfaces the day it stops holding.The ledger — 101 case arms, 2 path arms
carriedchangedsubsumedwithdrawnRounding the twelve non-carried arms up would be precisely CLOUD-908's recorded failure.
CLOUD-1218 — the lap journal charged a cold rebuild to the warm floor
This branch wedged its own landing loop on it, and it is the third container in two days to do so. A lap refuses on disk, the refusal says "free space outside ./target", the agent reclaims inside
target, and the rebuild that follows is recorded as a warm lap — teaching the warm ratchet a full rebuild's demand. Here that took the warm floor from a declared 6242MB to 24715MB against 8133MB free, after which every lap refused.basis_ofalready read the basis from the tree (since CLOUD-1157), but only over thedepsdirectories that exist — and the reclaim an agent actually performs removestarget/debug/deps, which drops that profile out of the walk and leaves a populatedtarget/release/depsreading warm..fingerprintis what makes the absence visible. Cargo writes one per profile it has built and leaves it behind whendepsgoes, so "is there a profile that has been built and now has nothing to build on" is answerable from the tree. Every.fingerprint's siblingdepsmust be populated. No.fingerprintanywhere falls back to the pre-existing reading, which keeps this a narrowing rather than a new requirement. It needs no guess about which profile the caller will build next — the guessbasis_ofrefuses and still refuses.The refusal also names
$GIT_DIR/batten-prune/laps.jsonwhen the floor in force is learned. That omission had a measured price: another container paid a full cold rebuild — 21519MB, ~20 minutes — reaching for therm -rf targetthe refusal does name.Two repairs the row specified were implemented and refuted, recorded in
basis_of's doc comment and on the row rather than left to be rediscovered:a_warm_laps_consumption_does_not_raise_the_cold_floorgoes red — it deliberately drives a warm lap consuming 22000MB against a 14000MB cold declaration, because that is the only shape distinguishing a per-basis ratchet from a shared one. The premise is also a unit error: the declared floor is a free-space budget, not a ceiling on consumption..any()to.all(). Does not reach a removed directory, for the reason above.The ratchet remains unbounded; that half is re-cut on the row as still open. CLOUD-1197 is affected — its version discriminator now has two basis-reading supersessions to cover, not one; flagged on that row while its PR is open.
The closing basis, and it is what blocked this PR from landing
Five consecutive
landlaps refused, and the last one is the diagnostic:verifysucceeded — a full cargo build, the whole cargo suite, all 2491 batscases — and
target-prunethen refused it at 14839MB free with no disk erroranywhere.
lap()computed one basis and used it for two questions: which ratchet bucketowns a lap's consumption, and which floor its remaining free space must clear.
The first is rightly the basis the lap opened under — what a lap cost is a
fact about the build that ran. The second is not: the floor asks whether there is
room for the build that comes next. A lap that opened on an empty
target/(correctly Cold) and closed on a fully built tree was still measured against the
cold floor — unreachable by construction, because building the tree is precisely
what spends the headroom a cold floor demands. Every
rm -rf targetrecovery anagent performs opens exactly that lap.
The tree's reading alone cannot decide it, and the first attempt at this repair
got that wrong.
basis_ofreadsdeps, so a lap opened Cold because an earlierescalation dropped
incrementalhas a fulldepsand reads Warm while its nextbuild really is a full one — the OR that function's own header describes, whose
second half the tree is blind to.
a_warm_laps_consumption_does_not_raise_the_cold_floorcaught it. So the discriminator is what the tree read at open, which
OpenLapnow carries beside the effective basis:
The field is optional, so a journal written before it parses rather than resetting
a clone's lap history.
A second defect, in this PR's own earlier
.fingerprintreadingWith the above landed,
verifystill refused its precondition as Cold on a treethat had just built cleanly.
basis_ofwalks the configured root unbounded, so itreaches every build tree nested inside it — and this repository's own suite
writes fixtures under
target/tmp/<case>/, one of which exists to model a profilewhose
depswas removed. That single fixture judged the whole repository cold.The predecessor had the same exposure and hid it: an
.any()over fixturedepsdirectories that happen to be populated read the tree warm, so the litter masked
rather than refused. Neither is a reading of this build. The walk is now filtered
to the two depths cargo writes —
<root>/<profile>/and<root>/<triple>/<profile>/.Measured after both:
warm floor 7264MBagainst12612MB free, a lap closingcleanly at 1183MB consumed, and
verifyreachingfast-forward-green.Shown able to fail in both directions, which is what the four new cases are
for: mutating the close back to
opened_basiskills two of them; mutating it toread the tree alone kills
a_warm_laps_consumption_does_not_raise_the_cold_floor.CLOUD-1216 — a glob cannot select a step for a deleted file
suite-bench-checkdecides set equality overgit ls-files 'tests/*.bats'in both directions, and its hk step was globbed on the corpus and the suites for exactly that reason. The glob could only ever deliver the first half. hk selects by matching changed paths, and a deleted path is not there to match — so the one commit shape the reverse direction exists for was the one shape that did not select the step. Measured on this PR: the commit deletingtests/derived-check.batspassed the full gate and left the corpus recording it, surviving four more commits.The step now runs unglobbed.
Shown able to fail through the real path, because the obvious probe does not discriminate.
hk check --planreports this step even for a commit touching onlycrates/**/*.rs, which neither glob matches — so it does not respect globs, and three selection probes against it "refuted" the row. I began correcting it before the control caught the mistake. The probe that decides is a commit:Cost measured rather than asserted: 531ms (531/530/531), replacing a comment that said "milliseconds". The wider class is surveyed on the row and
suite-bench-checkis its only member — 19 tasks readgit ls-files, but nearly all enumerate subjects, where a deletion removes a subject rather than creating a violation.A row this PR withdraws — CLOUD-1233
I filed it mid-PR claiming "the escalation drops
incrementaland the next build regenerates it, so the reclaim funds the next lap's overspend."batten.toml's own[[prune.regrowable]]comment refutes it, citing CLOUD-861's four-lap measurement:incrementalcarriescold = false— dropping rustc's incremental state does not force a full rebuild, and the laps following such an escalation consumed 228MB and 2852MB.incrementalwas not what it took.So there is no reclaim-then-regenerate cycle and the escalation is not implicated. I inferred a mechanism from two adjacent numbers instead of reading the config comment that already answered it, and filed a new row where a comment belonged.
The measurement survives and now lives on CLOUD-861, which owns the once-per-lap floor: consecutive laps at the same HEAD, differing only in
CARGO_INCREMENTAL, consumed 6859MB vs 613MB — 11x. Against a 7264MB warm floor, a lap that writes an incremental cache spends nearly the whole budget on it. That names what CLOUD-861 §1's exhausting phase actually is, and sharpens its ratchet: the same repository learns floors ~11x apart depending on an environment variable the journal does not record. It does not propose the lever —CARGO_INCREMENTAL=0on the verify path trades disk for CPU on every contributor's local gate, and would fail the acceptance I wrote on the withdrawn row, which required a lap with headroom to keep its cache.Five dead gates, all caught by the second tier
Writing this produced five predicates that loaded clean, passed their own
test_cases, and decided nothing:object.unionremoval fixturespatterns: &[]with input ascase fabricates the pattern data too, so deny and clean both passedline_sourcesmissing the workflowslinesobject the row would never buildThe last two are CLOUD-845's class arriving inside the test harness. A green
policy testis not evidence that a predicate can fire.The fifth is worse: the
ci-hygienepreset read three regexes from the[[pattern]]registry, which a preset can never resolve —policy.rssays so in as many words. Two rules gated nothing, and my own harness was supplying the patterns and hiding it. Literals inlined; the harness table is now&[], which is what a consumer hands a preset..claude/rules/policy-modules.mdtold authors to do exactly this and is corrected here — CLOUD-934 predicted the dead gate in those words.One regression, and what it taught
The preset broke
tests/prebuilt-lint.bats. That fixture copies the realbatten.tomlon purpose, so a rule ported from a task pointed at one directory into a config rule now judges every fixture tree inheriting the config. The fixture was right: a shipped ruleset must not refuse an ordinary minimal repository. The concurrency-group rule is scoped to triggers where two runs are two answers to one subject; every measured instance of the original defect is inside the narrowed set. Its ledger arm moved tochanged.Notes
CI_REQUIRED_CHECKS, andci.yml:716-731records it being removed from the fan-in because it never decided anything. Dispositioned in a comment below.config-deprecationsfails locally because the container has no git tags. It did; the tags were fetched and it reports 0 unannounced removals against v0.0.134. The underlying gap — a gate green in CI and red on every fresh clone — is CLOUD-1070's.closing-key-checkexpects a PR to close several — this PR names five and can record a claim for one.