fleet-floor's durable control plane — what the collector already does, and the two halves that are still missing #231
Replies: 8 comments
Correction — section 3 accused the machinery of something the label events refuteFound by re-reading What section 3 said
What #109's label events say
The door held, and it held three times faster than #212's 39 seconds. Triage opened it — a session applied The same correction lands on #236, which cited this as "the second instance found today of the same hole." It is not an instance at all: #128 was minted by triage, at What is actually true, stated smallerThe mechanism, read at the pin crew consumes (
So the residual property is real but narrow: Nothing is owed to heavy-duty/ceremony. I was about to open a discussion there, and the timeline is the reason I did not. The finding this thread should carry instead is one sentence about triage's own conduct: do not label a stray Nothing else in this thread changes. The reading of what already shipped in |
Correction — Q3 offers durability a seat on
|
0.1.1 |
#280 — a list of its own, cut from main tomorrow morning, 2026-08-03. Not #162, and not a polish gate: it is the acceptance-wave fixes that have already landed. |
0.1.2 |
#162 — the polish gate, retitled in the same ruling. It holds twenty open issues. |
0.2.0 |
#163, unchanged — and it is where this thread is already named. |
| #207 | back to its original scope, the 0.1.0 acceptance run; it no longer owns the 0.1.1 cut, which moved to #280 with the list. |
So the honest option set for Q3 today is #162 (0.1.2) or #163 (0.2.0). 0.1.1 is not an available answer: #280 takes whatever has landed by the cut and everything unlanded reverts to #162 by its own standing rule, so an issue minted from this thread today could not reach it. That is the third correction of this class in two days — #279 took the same one this morning, and the welcome thread took it in place — and it is worth making here because Q3 is the one question whose answer is a gate name.
The load-bearing half re-verified, because a stale problem statement is what put this thread here
This thread exists because #109's premise had gone stale under it — half the epic had shipped and nobody re-read it. So the claim this thread rests its own split on was re-checked rather than assumed, at a17b962, main's head at the time of writing:
"The snapshot is
self.snapshot, in memory, andfloor.pywrites nothing to disk — everyopen()in the file is a read."
Still exactly true. Five open() sites, all reads (the floor envelope, the roster, the probe, an agent conf, the index blob); the only writes in the file are sys.stdout and the HTTP response socket. No sqlite3, no json.dump, no write_text. Restart the collector and the fleet's observed history is still gone.
So the split this thread proposed is intact and both halves are still unbuilt — durability the small half, the action ledger the large one — and Q1 is still the cheap question that yields a buildable issue on its own.
One thing did move on the board, and it lands beside Q3 rather than on Q1 or Q2
#291 was minted 2026-08-02 at danmt's direction (blocked on #163): a droid is a chair with a CPU — one instance, one room, the room being that instance's details page. It is a read-path change to the same surface these projections would eventually feed, and its own open list asks whether its floor-only half justifies an 0.1.2 slice.
It answers neither Q1 nor Q2 — durability and an action ledger are orthogonal to how a page is composed — but it means the fleet-floor surface now has two candidate claims on the same window, and whoever answers Q3 should know that rather than discover it. Recorded, not absorbed.
Outcome unchanged: ask, waiting on @danmt, with Q1 still the cheap one. Only Q3's gate names move, and they move to #162 or #163.
Nothing is waiting on this thread. Re-checked across every open blocked body rather than inferred from the cross-reference list: no open issue's parsed blocker set names it, and #125, which cites it, declares Blocked by #123.
Triage, 2026-08-02.
Re-verified after four fleet-floor merges, and one inference to head off before someone makes itTwo things since yesterday's comment, one mechanical and one that arrives with this morning's corpus. The outcome does not move: ask, waiting on @danmt, Q1 still the cheap one. 1. The load-bearing claim, re-checked at
|
One word in my comment above is now false — the corpus is private, not public — and the claim it was arguing for is unchangedThis morning's comment called The argument is untouched, and the correction slightly sharpens it. The point was that the corpus is not evidence for Q1 — it is each box's For anyone following the citation: the public artifact is The load-bearing claim, re-checked at
|
Q3's option set is void — and the roadmap has since put a better answer on the board than either option this thread offeredMy comment at The better answer: this thread's subject is now a release, one arc wide#328 —
Its to-mint list carries the puller ("host-side, per-box, cadence + on-demand"), the store and its schema, and — the line that matters most here — "the floor and So Q3's honest option set today is #328 ( What that does to Q1 — it does not answer it, it re-shapes itQ1 asked whether fleet-floor's own snapshot durability is worth minting as one small issue, now, independent of the ledger. #328 answers a strictly larger question from the other side: a host-side store the floor reads instead of live boxes. Two stores, one of them per-process and one of them fleet-wide, is the duplication this thread was moved off the board for in the first place — "building this first would mean building a worse version of it." So the question to carry into And Q2 is untouched but no longer hypothetical. Does The load-bearing claim, re-checked at
|
Five fleet-floor merges later the load-bearing claim still holds — and Q3's two options are unchanged in membership and much closer in distanceFirst comment here since 2026-08-03 1. Re-verified at
|
The load-bearing claim holds at
|
The first of the three landings Q3's answer named is done — the record shape is on
|
| part | window | state |
|---|---|---|
| record shape | 0.1.3 — #327 |
landed. shared/bin/vitals.sh and shared/lib/common/tick-health.sh exist; the session record carries cost_usd= and sid=. #483, #484, #475, #485, #538 all CLOSED |
| store | 0.1.4 — #328 |
open, epic + release, window not opened |
| guarantee | 0.2.5 — #442 |
open, epic + release, window not opened |
The emission contract is worth re-checking rather than assuming, and it half holds
The boundary annotation made structured emission "a requirement on 0.1.3 rather than a request from 0.1.4" — versioned JSON on stdout for the probe, append-only JSON lines at a known box-local path for events — with the stated purpose that "0.1.4 ingests rather than parses" and the stated hazard that "a schema derived from formatting accidents propagates four windows forward."
The probe half is built. The append-only JSON-lines event path is the half I could not confirm: grepping the engine at 70bd7d4 for a JSON-lines emitter turns up shared/lib/duty-reaper.sh and nothing that looks like the events path the annotation describes — the session record is still a SESSION END line with key=value fields, now including cost_usd= and sid=, which is a formatted log line rather than a JSON object.
I am stating that as a measurement I took and not as a defect I am asserting. I did not read every consumer, and the annotation's requirement may be satisfied somewhere I did not look or may have been deliberately re-scoped inside the wave. But it is exactly the thing 0.1.4's init has to check before it builds an ingester, because the annotation's own hazard sentence is about precisely this: an ingester written against key=value log lines is the schema-from-formatting-accidents case, and it propagates forward from there.
Outcome: answered, and the next check is 0.1.4's init
Nothing is minted by this comment and no label moves. No ask for @danmt, and so still no return date.
Wake condition, re-stated so it is checkable rather than implicit: this thread wakes when 0.1.4's window opens. The one thing to carry into that init, so it is not re-derived: verify the append-only JSON-lines event path exists before designing the store against it — if what landed in 0.1.3 is a formatted log line rather than structured events, that is a 0.1.3 gap discovered at 0.1.4's door, and it is much cheaper to find it at the init than inside the ingester.
Triage, 2026-09-01.
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Why triage sent it back through this door
This is a good piece of design writing. The source-of-truth boundary section in particular — "SQLite is authoritative for accepted fleet-floor actions … the box remains authoritative for its actual runtime state" — is the sentence that stops a control plane becoming a fictional second fleet, and it is worth keeping whatever gets built. Four reasons it is a thread and not a board entry:
1. Half of it has already shipped, and nothing noticed
This is the load-bearing reason, and it is why minting children from the text as written would have been the expensive mistake.
The proposal's opening premise is that
fleet-floor"pays for live truth throughbox exec" on the page-read path: "every refresh fans out to boxes, results arrive slowly, one unavailable box stretches the request." That is not what the collector does today.floor.pyis already the background collector the proposal's scope section asks for:81687adFleet— "the telemetry snapshot, refreshed by one background thread";loop()polls on an interval and page reads serve the snapshotobserved_at)"generated, and the page renders staleness from it (test/stale.js)PROBE_TIMEOUT_S=45,PING_TIMEOUT_S=5,ACTION_WORKERS=8, a single-flight_poll_lockthat coalesces click bursts, and a separate ping tier so the fast signal is not queued behind a 45s evidence probetest/concurrent-actions.pyis the #44 guard, asserting peak concurrency across a wedged fleetSo acceptance criteria 1 and 3 — "loading the ordinary fleet dashboard performs no synchronous fan-out" and "a slow or unreachable box does not block the rest of the fleet from being displayed" — are already met by merged work. An epic whose first two criteria are green before anyone starts is an epic that will send its first builder to rebuild something that exists. This board has the lesson written down already, on #116: "a
Blocked bydeclaration is a snapshot, and nothing re-validates it." A problem statement is a snapshot too.2. What is actually left is real, unmet, and a different shape than the title
Stripping out what shipped leaves two genuinely open gaps, and they are not one piece of work:
self.snapshot, in memory, andfloor.pywrites nothing to disk — everyopen()in the file is a read. Restart the collector and the fleet's entire observed history is gone; there is no last-known-good to show during an outage, and no separate record of the most recent failure. This is the proposal's "persist each box's last successful observation and its most recent failure separately", and it is the small half.The first is small and buildable now. The second is not, and its cost is mostly in
crew floorgrowing a second long-lived process — see Q2.3. It organizes nothing, while claiming to
It has carried the
epiclabel since 2026-07-28. Under LABELS.md an epic "organizes other issues via a dependency-ordered task list". Four days on it has zero children; its "Road" is a numbered list of work nobody has minted. That label is not true.And — this is the part worth recording — theepiclabel is why nobody caught it: the work-queue sweep's invariant exemptsepicfrom the one-of-ready/claimed/blocked/post-mergecheck, so it never drew aneeds-triageflag. #212 arrived on 2026-07-31 and was flagged in 39 seconds. #109 arrived on 2026-07-28 and was never flagged at all. Anepiclabel on an untriaged issue is a hole in the door.4. Its own text defers the cut
That review is this thread. The proposal is explicit that it is not yet a set of work orders, and triage agrees with it.
The questions whose answers let triage mint
Three, in the order their answers unblock work. @danmt — Q1 is the cheap one and yields a buildable issue on its own.
Q1 — durability alone, now?
Given that the collector exists and only its memory is volatile, the small half stands on its own: persist the snapshot so it survives a restart, keep last-success and last-failure apart, and serve a last-known-good with its
observed_atduring an outage. No action ledger, no locks, no worker, no lifecycle model.sqlite3is in the Python standard library, so this keeps the stdlib-only propertycrew flooradvertises andREADME.mdrestates — that constraint is not a reason to avoid SQLite, and a flat JSON file would be the weaker choice for the same cost. Is that worth minting as one issue today, independent of everything below?Q2 — does
crew floorbecome a supervised service?This is the question that sizes the large half, and the proposal does not name it.
Today
crew flooris one process an operator runs. A durable action ledger needs a worker that outlives an HTTP request, and the proposal's recovery semantics ("worker/browser/API restarts do not lose accepted actions") only mean something if something restarts the worker. Two answers, very different bills:floor.py— single writer, no new process, dies and restarts with the page server. Cheap, and honest as long as the docs say a stopped collector means stopped actions.Which one is being asked for? The answer decides whether this is a fleet-floor change or a crew-platform change.
Q3 — which gate? (danmt's alone)
#163 (
0.2.0) names #109 under "the console must show what a human needs", beside #125, #126 and #129. But a control plane is not a console feature — it is fleet-floor's read and write path, and the large half touches every mutation the console has.0.1.0is in its acceptance phase (#207) and0.1.1(#162) is the polish gate.Prioritisation is yours. Q1's answer may make this moot for the small half — durability is cheap enough to ride
0.1.1— while the ledger waits for a gate with room.Where this sits on the board meanwhile
#163 already names it on the
0.2.0road, and that reference stands — it now points here. #126 (a box's OS, health and process view) declaredBlocked by #109; that declaration is being re-pointed, since it was waiting on the collector substrate that turns out to already exist.Nothing on the board is blocked by this thread. #207, #162 and #163 are unaffected, and every
readyissue continues.The proposal, as filed on #109 — verbatim
Problem
fleet-floorcurrently pays for live truth throughbox exec.That is tolerable for an operator command, but a poor read path for an application: every refresh fans out to boxes, results arrive slowly, one unavailable box stretches the request, and the UI has little durable state to explain the wait. Mutations have the sharper version of the same problem. An operator can click
on,off, hire, or another conflicting action repeatedly while an earlierbox execis still running. The browser knows that it sent an HTTP request, but the system does not expose a durable answer to:This is a correctness and operability gap, not only a loading-state problem.
Direction
Introduce a deliberately small fleet control plane backed by SQLite.
box execremains the mechanism for actions, installation, recovery, and detailed live diagnostics. It stops being the synchronous source for routine page reads. The application reads projections from an API backed by SQLite and reports their freshness.A click creates a durable action and returns an action ID promptly. A worker performs the slow command, while the UI follows recorded state rather than holding one opaque request open.
Scope
1. SQLite store and migrations
The first useful model should stay small:
boxes/ snapshots: last observed state, observation time, probe result, and error/freshness metadata.actions: type, target, requester, idempotency key, lifecycle state, timestamps, timeout, result, and bounded output/error detail.locks: resource, owning action, expiry, and fencing/version token.Names and normalization belong to the implementation issues; the invariant is that observations, commands, and their status survive an HTTP request or browser session.
2. Background collector and projections
observed_at) and distinguishes current, stale, unknown, and unreachable.3. Durable action ledger and worker
runningaction are explicit: retry only when safe, otherwise reconcile and surface an indeterminate/failed result.4. Conflict locks
Locks are scoped to the affected resource, not a global “disable the whole app” switch:
on,off, restart/delete if present) conflict per box;The API enforces conflicts. Disabled buttons are feedback, not the safety boundary.
5. Action-aware UI
Source-of-truth boundary
box execaction results are evidence, not a substitute for subsequent state reconciliation.This boundary keeps the database useful without allowing it to become a fictional second fleet.
Road
Follow-on issues should be cut from this epic after the boundaries are reviewed. A likely order is:
box execfan-out.Each step must remain deployable and testable; this epic does not authorize a big-bang rewrite.
Acceptance criteria
on/offconflict, worker death, stale locks, late completion, a wedged box, and stale snapshot rendering.Non-goals
box exec; it remains the command and diagnostic channel.Later: centralize the GitHub transport
Once the SQLite collector, durable action model, locks, and projections are operating successfully, revisit crew's GitHub integration.
The same control-plane foundation could support one central GitHub collector that normalizes repository events into a durable event log and gives duties pull-based inboxes, acknowledgements, retries, leases, and replay. That may reduce duplicated polling, rate-limit pressure, credential variance, and coordination races across agents.
That is intentionally follow-on work, not hidden scope here. GitHub should remain authoritative for repository facts; a future transport would be authoritative for delivery, acknowledgement, leases, and cached projections. Centralize observation first and assess measured pain before deciding whether GitHub mutations should also move behind it.
All reactions