A sentinel — an ops role that reads the fleet's own telemetry and flags what a human should see #236
Replies: 9 comments
Correction — the "second instance of the same hole" was not an instanceStruck in the body above, against both issues' label events rather than their threads. The full working is on #231; the short version:
So there is no defect in the sweep, and nothing is owed to heavy-duty/ceremony — I was about to open a thread there and the timeline is why I did not. What still stands, unchanged: the paragraph above the struck one — |
Correction — this thread's header says #128 closed, and since
|
The report that answers Q1 is claimed and in flightShort update, from label events rather than from this thread. My comment above says #290 is Q1's wake condition is unchanged and is now days rather than a gate: the indicator set that report ends up using is this thread's answer to which indicators, and what threshold makes one worth waking a human over, derived from two real working days instead of guessed. Triage mints the summariser from it. Q2's sub-question stays here for @danmt and is untouched by any of this: does a sentinel claim carry the evidence rows it was drawn from, so a human can check a finding without re-deriving it? The structural half — the sentinel flags, it never acts — is answered by #295's chair table and needs nothing further. #128 remains Triage, 2026-08-02. |
Q1 is answerable from data as of this morning — and the premise underneath it ("there is nothing to summarise yet") turns out to be falseThis thread's outcome has stood on one sentence: "there is nothing to summarise yet — #121's ledger (#122) and its aggregation (#123) do not exist, so a sentinel has no input." Two working days of fleet telemetry became public this morning ( What the fleet's own
|
Q1's wake condition fired — the report is on
|
Both halves of the substrate left the board at
|
| #128 — the sentinel charter | closed 2026-08-03T15:54:24Z in the operator-directed board-hygiene pass (close comment), carried by epic #163 (0.2.0). blocked is still on it — a close does not move the queue label — so the label a scanner reads is stale and the close event is the fact. #295's checklist records it, with the same carrier. |
| #123 — the aggregation | closed 15:54:17Z, carried by #328 (0.1.4 — collector + database), where it is named as "the collector's first consumer". |
| #122 — the ledger | closed 15:54:15Z, carried by #327 (0.1.3 — telemetry). |
So the sentence this thread's outcome rested on for four days — "there is nothing to summarise yet; #122 and #123 do not exist" — has changed shape twice this week. This morning the corpus falsified it as a fact (the fleet's own duty.log was already summarisable). This afternoon the two issues it named stopped being open work at all.
The indicators are scheduled, and they are scheduled before the chair
This is the part worth having on the record. #327 (0.1.3 — telemetry) carries, in its to-mint list:
"Tick-health surface: last-tick age, skip/hold counts, outcome streaks — the numbers the 2026-08-01/02 report mined by hand."
That is the raw material of three of the four indicators I recommended this morning — 1 (a session that did not finish → outcome streaks), 4 in its hold-age form, and 3's per-kind baseline — landing in 0.1.3, while the chair that reads them (#128, carried by #163) is 0.2.0. Indicator 2's signal already ships: #314 went out in 0.1.1 and writes "one WARN in duty.log naming the PR, the head and the count" on the trip.
That ordering is danmt's 2026-07-28 ruling coming true by the roadmap rather than by argument — "real programming to turn telemetry into summaries with multiple indicators; the sentinel agent simply reads the summary." The deterministic half is two arcs ahead of the model half, which is the shape this thread called load-bearing.
What is scheduled is the numbers, not the judgement. #327's bullet names surfaces; it names no threshold, no digest, and nothing about what is worth waking a human over. That is still Q1, and Q1 is still @danmt's.
One thing to carry into 0.1.3's init rather than discover in it, because it is this thread's own correction and the epic's wording does not have it: "skip counts" is the form the merged report declines to call a signal — 568 busy ticks in two days, 26.1%, which is "the non-blocking lock declining duplicate work" and a baseline rather than a finding. The signal is the field beside it: hold age, observed at 253 s to 4,800 s. An indicator built from the skip count cries wolf 568 times a fortnight; one built from hold age fires on the 4,800-second hold. Recorded on #327 as well, so a builder meets it where the work is.
Outcome unchanged: ask, waiting on @danmt
Q1 — which indicators, and what threshold makes one worth waking a human over — is now a pick from a measured list with a release window under it, and it decides what the sentinel is on its first day: a fleet-health watch rather than a spend watch. Cost indicators join when the ledger returns with #327's window; they are a second slice, not a precondition. Q2 — does a sentinel's claim carry the evidence rows it drew from — is untouched.
Both decided constraints are untouched: the summariser is deterministic code, and the sentinel flags, it never acts.
Nothing is stalled by the ask, and it is now stalled less than it was: #128 is closed, so nobody can pick it up at all, and re-opening it is triage's act at 0.2.0's init. Nothing is waiting on this thread — re-derived at this write with blockers.jq's clause parse over all 22 open blocked bodies: every one names #346, none names this thread.
Triage, 2026-08-03.
The substrate window is now the next one — and the report's largest unpriced class got a measured instance, an engine detector, and a new
|
Q1's first half got a taxonomy this morning; its second half is untouched — and the same session handed this chair two detection blind spots it should be built knowing aboutOne delta since my comment of 2026-08-07 1. What the design session did and did not answerThe reading taxonomy and the boundary annotation, both 2026-08-09, decide which readings exist, whose store holds them, and how long they are kept. #327's to-mint list carries the sentinel's own primary input verbatim: "Tick-health surface: last-tick age, skip/hold counts, outcome streaks." No threshold was ruled anywhere in either note. So Q1 splits cleanly now: its which is largely answered by classification, and its what number wakes a human is exactly as open as it was, still due at the same init, and still the half that decides what this chair is on its first day rather than what it can see. 2. Two blind spots this chair inherits, both recorded as knowingly acceptedNeither is a defect and both are invisible from the sentinel's own side, which is why they belong on this thread rather than on the epic that accepted them.
3. OutcomeAnswered, still held. No question here is resolved and none is re-raised; Q1 goes to #327's init as before. What is new is that the which half arrives pre-classified, and that two things this chair will not be able to see are now written down before it is specified rather than discovered by it. Triage 2026-08-09 |
The init this thread's Q1 was placed at came and went, and the question was not put — recorded, and re-placed so it cannot go dark twiceTriage 2026-09-01, re-measured at This thread carries no return date, correctly — its wake is an event. The event fired and nobody was watching. No sweep, no nudge and not the What was placed, and what happened to the placementMy comment of 2026-08-09 placed Q1 — which indicators, and what threshold makes one worth waking a human over — at "the release-init of the window after the one being cut", noting Every step of that has since happened:
Q1 was not put at that init. It is not on #327, it is not on this page, and no The which half answered itself; the threshold half is exactly as open as it wasThat comment had already split Q1: its which was largely settled by the reading taxonomy, and "what number wakes a human" was "still the half that decides what this chair is on its first day rather than what it can see." The sentinel's primary input is now built, not specified. Measured at
No threshold was ruled by any of it. I re-read the landed work rather than the annotations describing it: these ship readings and surfaces, and not one names a number at which a human is woken. So the split holds exactly as stated — the which is now answered by working code rather than by classification, and the what number is untouched, thirteen days after the init it was due at. The Outcome: ask, re-placed at
|
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Why this is a thread and not an issue
Its own header, quoted exactly:
An issue whose central design question is unanswered is not a work order; a builder handed it would guess, and guessing is the failure the triage door exists to prevent (TRIAGE.md, which names the move directly: "Mint an issue to 'discuss' something — that is a discussion").
The
epiclabel was false, and it hid the problemUnder LABELS.md,
epicmeans "organizes other issues via a dependency-ordered task list". This one organized zero issues across four days. It had no children, no checklist, and no dependency order — it was a design record wearing an organizer's label.Worse, that label put it in a contradictory state on the board:
epicexempts an issue from the work-queue invariant, so it carried no queue label at all, while its Dependencies section said in prose "Blocked by #123." The board said epic; the body said blocked. Both cannot be true, and a scanning agent reads the label.This is the second instance found today of the same hole — #231 (ex-#109) was a stray that went four days without aneeds-triageflag for exactly this reason, because the sweep's invariant exemptsepic. Anepiclabel is a place board state goes to hide. Worth its own fix; recorded here and on #231.What survives, because it is the best thing in it
Two decided constraints, neither of which is an open question:
1. The sentinel flags. It never acts.
That is the fleet's own spine — detection and judgment must live in different places — correctly re-stated for the case where the model is the detector. It is the single most important line in the record and it must survive into anything built.
2. danmt's ruling of 2026-07-28, which settles the architecture:
Most of the work is not an agent. It is the summariser — deterministic code over the ledger — and the agent is a thin, optional layer on top of it.
The outcome: ask — and the substrate comes first regardless
The honest position: there is nothing to summarise yet. #121's ledger (#122) and its aggregation (#123) do not exist, so a sentinel has no input, and the record's own awkward observation stands — "a sentinel costs tokens to tell you about token costs" — which must be measured against #121's numbers rather than assumed.
Two questions, for whenever the substrate lands:
@danmt — (1) is answerable before #122 lands and is the cheaper half by a wide margin. Answer it and triage can mint the summariser as a normal issue behind #123, with the agent deferred to a second one.
Board effects. PR #230 references this as
Refs #128for the ARGUS prototype; that reference now points here. The prototypes are design records too, and unaffected.Nothing is waiting on this thread.
The design record, as filed on #128 — verbatim
Why it is the right instinct
Everything the fleet is about to start collecting — token cost per session and pool (#121), whatever telemetry #126 adds — has the same problem the Actions bill had: it is only useful if somebody looks. The 2026-07-25 Fable burn was diagnosed by hand, after the fact, because nothing was watching. Building the instrument and then relying on an operator to read it re-creates the gap one layer up.
A sentinel is the fleet doing to itself what it already does to repositories: sweep, notice, and wake a human.
The design constraint this fleet already settled
From a box's own self-assessment, and it is the spine of every duty module:
A sentinel inverts the usual arrangement: here the model is the detector. So the rule has to be re-stated for it, and it is the strongest thing this record can contribute:
The sentinel flags. It never acts. No label writes, no configuration changes, no stopping a box. It produces a claim, addressed to a human, with the evidence attached. Anything else makes a hallucinated finding into a fleet action.
The awkward part, stated plainly
A sentinel costs tokens to tell you about token costs. That is not disqualifying — one summarising session per day against five boxes' worth of ledger is cheap next to what it watches — but it must be measured against #121's own numbers rather than assumed, and it is exactly the sort of thing #120 exists to be sceptical about.
Ruled by danmt, 2026-07-28, and it settles the architecture: "I literally plan to do telemetry and use real programming to turn that into summaries with multiple indicators. The sentinel agent simply reads the summary and gives an assessment extra."
That is the same shape #91 found, applied here: most of the work is not a model's. The engine can compute the deltas — this pool spent 3× its trailing average, this box has not ticked in two boundaries, this session ran to timeout four times today. The model's job is the last mile: reading a compact digest and saying which of those matters and why. Feeding it raw logs would be the expensive way to get a worse answer.
It is a new role, and that has a known cost
crew's roles are triage, builder, reviewer. A sentinel is a fourth, and #71 is explicit that a new role changes the engine and its duty lifecycle — a profile is configuration, a role is engine work. #71 also says not to build a plugin system in anticipation, and to generalise only after two more roles exist.So this is expected to be somewhat messy to add, and that is accepted rather than a reason to defer. Worth knowing before someone estimates it as "another duty module".
The three layers
danmt: "it's a lot of work, probably spans multiple issues." It is, and they are not one deliverable — three layers with a real dependency order and different skills.
Follow-on issues should be cut from this epic after the boundaries are reviewed, the way #109 does it. A likely order:
Layer 1 is the critical path and is useful alone — indicators with no assessment are still a dashboard. Layer 2 without layer 1 is a model reading logs, which is the thing being avoided.
Open questions
notify.sh's Telegram channel already exists for operator escalation, and reuse beats a second path.Dependencies
Blocked by #123 — there is nothing to summarise until the ledger exists and something aggregates it. #122 and #123 are the substrate.
Benefits from #109: trend detection needs history, and "this pool spent 3× its trailing average" is a projection over stored snapshots rather than a scan of log files. Benefits from #126 for the same reason if system telemetry lands.
Related to #71 (a role is engine work) and #120 (be sceptical of anything that spends tokens to save tokens).
All reactions