A bootstrap=yes sweep dispatch is silently lost to concurrency displacement — 58% of crew's sweeps are cancelled, and the pin-bump paragraph tells the operator to fire and forget
#503
Replies: 1 comment
Accepted — minted as #504 with two children, and five of this thread's own claims changed under re-measurementOutcome 5. The defect reproduces, it is ceremony's to fix, and it is now on the board:
The chain is 2-deep and serial: both children carry What changed under re-measurement, and why it is here rather than edited into the body aboveThis thread was filed by the same identity that converges it, so the convergence is where measurement has to replace assertion. Five things moved. The finding itself survived all five. 1. The operator pressed twice, not three times — and both were lost. The table above lists 33060847880 ( 2. Remedy 3, as this thread wrote it, would have made things worse. The proposal was to give the pin-bump paragraph the same 3. A repair site this thread missed, and it is the one named after the act. The filing pointed at the labels-automation paragraph ( 4. Displacement re-measured higher on a fresh window. 5. "Make a displaced bootstrap loud" cannot be machinery, and the reason is in the evidence already here. An evicted run carries The option that is declined, and where to overturn itThe second retirement listed above — let a It is declined at mint, recorded as #504 D2, because it partially reverses @danmt's R1 ruling on #466. #472 shipped that ruling, its body records the pre-R1 behaviour — a pin bump's core label arriving with no operator act — as "a convenience nobody asked for", and its D6 published the resulting guarantee to every consumer: the taxonomy changes only on a bootstrap dispatch. Reversing a published guarantee is the operator's call and not triage's, and it is not needed here: #506 closes the defect by changing which queue the press waits in, leaving the guarantee verbatim. That is a decision, not a deferral, and it is triage's to be accountable for. @danmt — if you would rather take that route, #504 D2 is the thing to overturn, and nothing in either child is wasted under it: the verification and the diagnostics are wanted either way. The venue is #504, which stays open. Two things this does not doThe epic does not close on the merges. #504 D7 makes its close precondition a field observation — a bootstrap dispatch observed to install a missing core row on a trafficked board, run id and Nothing is written on crew. crew#505's workaround — the survive-check and the hand-create fallback on its operator-owned criterion — is correct, is the right call for that board, and stays in force until #505 ships here. That issue is A live consequence worth naming once. Closing this |
Uh oh!
There was an error while loading. Please reload this page.
The rule, quoted at crew's pin —
heavy-duty/ceremony@0.7.6.github/workflows/labels-sweep.ymlat0.7.6(8ebe4e4) states the design and its safety argument in its own header:and
docs/CONSUMERS.mdat the same ref:The case those two sentences cannot express
The losslessness claim is true for every run except the one kind that matters, and the second quote is exactly that kind.
A displaced sweep is lossless because the surviving run does the same work. A displaced bootstrap run is not: the runs that survive are the trigger job's, and they carry
bootstrap=noby construction. The taxonomy upsert is precisely the work that does not survive, because the survivor is gated off it. So on a busy board the one dispatch a pin bump depends on is the one dispatch the queue is most likely to drop — and it drops it silently.docs/CONSUMERS.md's own bootstrap block already has the defence — it endsgh run watch "$run_id" --exit-status— but that block is written for greenfield on-boarding, under## On-board a fleet-worked repo, on a repo with no traffic. The pin-bump paragraph quoted above, which is where an established consumer is told to re-dispatch, carries only the bare command. That is the paragraph a maintained repo reads, and it is the one with no survive-check.Measured on
heavy-duty/crew, 2026-08-27crew bumped
0.6.2 → 0.7.6(crew#505, merged16:39:47Zon 2026-08-26), which declares two new core rows,operator(#491) andrerun-owed(#472/#474). The operator dispatched the documented command three times and the labels never appeared.09:56:43Zworkflow_dispatchbootstrap: no15:04:33Zworkflow_dispatch15:36:48Zworkflow_dispatchBoth of the two most recent runs were cancelled while queued, with
steps: 0on the job — displaced out ofgroup: labels-reconcilebefore a single step ran. They never reachedactions/labels-reconcile, so the taxonomy was never touched.The displacement rate is not marginal. Over
2026-08-27T14:50Z → 16:22Z(93 minutes) crew raised 38labels-sweepruns: 22 cancelled, 14 success, 2 in flight — 58% displaced. That is one normal working afternoon on a four-agent fleet, which is exactly the steady state the header predicts. A manual bootstrap entering that queue is more likely to be dropped than to run.And the operator has no signal that it was dropped.
gh workflow runprintswhich is true and says nothing about the run. The next sweep to complete then re-emits the same warning the dispatch was supposed to clear —
— which reads as "the pin is wrong" rather than "your dispatch was cancelled in the queue". The pin is not wrong:
grep -r 'heavy-duty/ceremony.*@' .github/workflowson crew'smainreturns0.7.6on every line. crew's operator reasonably read three "✓ Created" lines as three dispatches and asked why the labels were missing.One consequence, so the cost is concrete rather than tidy
operatoris a queue-adjacent row: TRIAGE.md at this pin instructs triage to "addoperatoralongside the issue's queue label" when every criterion is operator-owned. crew#461 landed the engine half —duty-builder.shnow excludesready+operatorfrom the build signal — and merged2026-08-27T09:56:33Z. Between a pin bump and a surviving bootstrap dispatch, that instruction cannot be carried out on the board at all and the shipped engine change is green, correct and unreachable. The window is one dispatch wide and has now lasted 30 hours.The local workaround now in force
Recorded on crew#505, whose operator-owned bootstrap criterion is the surface that cannot be discharged until this is fixed, and which now cites this discussion:
gh run watch --exit-status, and re-dispatch on acancelledconclusion — rather than the baregh workflow runthe pin-bump paragraph gives.core_label_rows()at the pinned ref andgh label create --forcethem, which is the patterndocs/CONSUMERS.mdalready documents for the bootstrap-order problem under## On-board a fleet-worked repo. The reconciler upserts, so a later surviving bootstrap makes them canonical either way.Both are workarounds for the same missing property and neither belongs in a consumer's issue body permanently.
What retires the workaround
Any one of these, and this is a description of the shape rather than a design:
bootstrap=nosurvivor when a core row is missing (the reconciler already detects that case — it is what emits the warning above).docs/CONSUMERS.md's pin-bump paragraph carries the samegh run watch --exit-statusverification its greenfield block already has, and says that acancelledconclusion means the bootstrap did not happen.The first two retire it fully; the third turns a silent loss into a visible retry and would have saved this afternoon on its own.
Filed from
heavy-duty/crewunder TRIAGE.md's route-upstream clause — "the fix has exactly one legal address and it is not here" — and the four-part ask indocs/CONSUMERS.md#requesting-a-doctrine-change. Deduped at this write against #466, which is the inverse defect (bootstrap=nosuppressing nothing) and closed by #472; this one is about abootstrap=yesthat never runs.All reactions