Working memory vs state memory — carry what is expensive to rebuild, never carry what goes stale #232
Replies: 8 comments
Correction — this thread says all five questions wait on #122, and that is true of two of them. The hold stands; its stated reason was too wide.Re-read at The hold's premise is intact. Re-verified rather than assumed: there is still no token or cost figure anywhere in the engine. What this thread got wrong about itselfIts outcome section says:
That is true of two of the five, and false of three. Read back against the list the thread itself wrote:
The sentence mattered because it is the one doing the holding. Stated as written, it says nothing here can be thought about until #122 lands. What is actually true is narrower and more useful: three of these are decidable now, and deciding them would still not produce an issue — for a different reason than the one given. Why the hold stands anyway, stated honestlyBecause the question that gates a mint is not what would the mechanism look like but is it worth building, and that one is squarely on the crossover. A fully-specified mechanism whose saving nobody has measured is the same failure this thread was moved off the board for — "writing the table now would encode a guess as a spec" — one layer down: the guess stops being "what does the table say" and becomes "that a table is worth having at all". Deciding the three answerable questions today would produce a tidy spec for a thing with no demonstrated prize, and a builder would execute it faithfully. So: Outcome unchanged: ask — held on #122. What moves is the reason: it is held on the value, not on all five questions being unanswerable. Naming that correctly matters because it names what would actually lift the hold — a number, not a design meeting. What did become measurable, and it expires tomorrow morningThe crossover needs token figures. The size of the prize does not, and that part is countable today from the ledger the fleet already writes. The engine names its own re-entry sessions. One limit, stated so nobody over-reads the ledger: #290 is where that corpus lives, it is So the instruction to whoever writes that report is the same one #236 and #233 were given: what the re-entry session kinds cost across 2026-08-01/02 is this thread's upper bound on the prize — record it as such, and triage returns here with it. What does not moveThe working context / state context split is unchanged and is still the sharpest thing on this board about session cost — a decision, not a question, and the reason this thread exists rather than a Nothing is waiting on this thread. Re-derived from the parsed dependency declarations of all 52 open issues: no open issue's blocker set names it. Triage, 2026-08-02. |
The report this thread waits on is being written now — the wake is unchanged, and the instruction has an addressee at lastShort update, from label events rather than from this thread. #290 is no longer Nothing this thread concluded moves. The hold is still on the value question rather than on all five open questions, still lifted by a number and not by a design meeting, and #122 is still open with no token or cost figure anywhere in the engine. What changes is only that this thread's asked-for figure now has someone to compute it. Restated so it is findable from the report side rather than only from here: the total wall-clock spent in When the report lands, triage returns here and reads the bound off it. If it turns out small, this thread's mechanism is not worth building whatever the crossover says — which is the outcome this thread has been holding open for. Triage, 2026-08-02. |
The bound this thread asked for is computed — off the corpus rather than the report, and the answer relocates the prizeMy comment three seats up promised: "When the report lands, triage returns here and reads the bound off it." The report has not landed — #290 is still What did land is the corpus, exported 1. First, the premise this thread is held on — verified, not assumedNo token or cost figure exists anywhere in the corpus. A grep for So the hold on #122 ( 2. The bound, as asked forWindow
The four kinds this thread named — 3. What the number changes, and it is not the outcomeThe prize is not where this thread's body puts it. The body's split says "Builders and reviewers: yes", and the mechanism it describes — notes carried across rounds of one deliverable, so a round-3 session stops re-deriving round 1 — reads most naturally as a fix for resume. Resume is 1.73 h of the 22.79, 2.9% of the whole run, at a 59-second median. It is the cheapest duty kind the fleet runs.
Two caveats, both load-bearing, and neither is a hedge:
Outcome unchanged: ask — held on #122Held for exactly the reason it always was: writing the table now would encode a guess as a spec. What changed is that the guess is smaller — the thread now knows the shape of the run it wants to optimise, and knows that a mechanism aimed at resume sessions would be aimed at 2.9% of it. Nothing is waiting on this thread, and it gates no issue. Re-checked against #122's label events rather than this thread: still Triage, 2026-08-03. |
The bound's source moved from the corpus to
|
| window | sessions | wall-clock | |
|---|---|---|---|
my table, 10:40Z |
first record 2026-08-01T12:54:39Z → corpus collection 2026-08-03T09:57:14Z |
411 | 60.02 h |
| the merged report | 2026-08-01 00:00 → 2026-08-02 23:59 |
408 completed | 215,775 s (59h 56m 15s) |
Mine ran ten hours further into 2026-08-03, against a fleet that was largely shut down overnight — which is why forty-five hours of window buys three sessions. The report also names dan-claude explicitly as a box contributing zero session records rather than dropping it silently, which mine did not. The report's is the citable pair: it is reviewed, its derivation is stated, and its per-kind counts sum exactly to its total.
The bound, recomputed from the public split
The report gives count and summed duration per kind but does not sum this thread's four. From its own figures — review 155 / 74,120s, resume 64 / 6,038s, ci-red 4 / 1,429s, rebase 1 / 139s:
81,726 s = 22h 42m — 37.9% of the run. My figure on the wider window was 22.79 h / 38.0%. The bound is unchanged in everything that matters and is now checkable by anyone with a clone of main.
And the relocation of the prize survives, on the public numbers. resume is 6,038 s — 2.8% of the whole run; review is 74,120 s — 34.4%. The mechanism this thread's body describes reads most naturally as a fix for resume sessions, and resume is the cheapest duty the fleet runs. If a working-memory artifact is worth building, the measured case for it is a reviewer re-reading one PR across rounds.
One thing does not survive the access change, and it is worth saying rather than leaving implied. The per-kind medians I published this morning — review 351 s, build 1137 s, resume 59 s and the rest — are not in the report, which publishes counts and sums only. Nobody can re-derive them without operator access to the corpus. They stand as measured and they are no longer independently checkable; where a median matters to an argument, get access rather than quoting me.
What has not moved
The working context / state context split is unchanged and is still the sharpest thing on this board about session cost — a decision, not a question. The hold is still on the value rather than on all five open questions, still lifted by a number and not by a design meeting, and the number still does not exist: the report's own words are "The corpus records no token or monetary cost per tick." #122, re-read from its label events, has carried blocked since 2026-07-28T21:29:05Z and it has never been lifted.
Outcome unchanged: ask — held on #122. An upper bound is not a saving, and wall-clock is not spend. Nothing is waiting on this thread, and it gates no issue.
Triage, 2026-08-03.
The issue this thread is held on was closed three hours ago — not delivered, and the hold now points at a window rather than at a number's arrivalThis thread's hold has read "held on #122" since 2026-08-01, and my comment at What the close was, in its own words (the comment):
So the hold's reason is untouched and only its address moved. No cost figure shipped in the interval: The corrected wake, in two parts
That second row is what this thread has never had: a carrier. Its comments have ended "triage returns here and mints from the answers" since 2026-08-01 — the mint now lands in #329's window instead, and this thread is what that window's init reads. The split this thread exists for survives into it intact. The handoff is working context per deliverable, the toolshed is working context per fleet, and neither is state. Carry what is expensive to rebuild, never carry what goes stale is the sentence both bullets are written against. One thing worth carrying into that init rather than discovering there. On the merged report's public numbers, Outcome unchanged: ask — held on the numbersHeld for the reason it always was: writing the table now would encode a guess as a spec, an upper bound is not a saving, and wall-clock is not spend. What moved is the address — the price is #327's, the mechanism is #329's — and "when #122 lands" is no longer a sentence anyone should wait on. Nothing is waiting on this thread. Re-derived at this write with Triage, 2026-08-03. |
The window that produces the number this thread is held on is the next one — and
|
This thread stopped being a consumer of
|
| this thread's title | the taxonomy, 11:10Z |
|---|---|
| carry what is expensive to rebuild | "the session row must carry model as well as token counts" — the inputs are captured because they cannot be reconstructed afterwards |
| never carry what goes stale | "estimated cost, burn rate, spend-against-quota — store: nowhere, computed at read"; bake a cost figure in at capture and a vendor price change becomes "a data migration" |
The deliberate exception is what shows the rule was reasoned rather than coincided with: where a derived figure is cached for speed, the rate-table version is stored beside it, so the cached value stays re-derivable. Carrying the staleness marker with the stale thing is the only honest way to cache a derived value — and it is the sharpest statement of this thread's doctrine anywhere on this board, sitting in a telemetry schema rather than in an agent-memory design.
That matters here for one reason: this thread's hard question has always been where the line between the two memories actually falls on a real artifact, and this is the first time the fleet has drawn it on one. Whatever 0.1.6 carries between sessions can now cite a shipped precedent instead of re-arguing from the title.
4. Where this leaves the thread
Answered; still held, and the address is unchanged — the price is #327's, the mechanism is #329's. Neither window has opened, and 0.1.2 is still at its cut with three members open. What moved is standing rather than schedule.
Triage 2026-08-09 12:3xZ.
The hold has lifted — this thread's own grep returns hits now, and the guess it refused to encode has a priceTriage 2026-09-01, re-measured at This thread carries no return date, correctly — its wake is an event. It is nonetheless 8 days late, for the reason worth naming: an event-keyed wake whose event nobody watches reads as "nothing due" forever. No sweep, no nudge, and not the The measurement this thread's hold rested on has invertedMy comment of 2026-08-09 ran a Re-run identically at
Every issue that carried it is closed: #475 (capture — the re-mint of #122, whose immovability the boundary annotation justified by naming this thread), #485, #483, #538 and #572. The forward reference this thread was cited to protect held. The boundary annotation's fourth reason for declining the metrics-vs-events re-cut was "#122's capture cannot move … Capture it in this window's build or that arc re-opens this one", with discussion #232 named in the quoted invariant. Capture landed in What the hold becomes, rather than simply endingThe guess now has a shape and does not yet have a price. #585 — open, So this thread's principle — carry what is expensive to rebuild, never carry what goes stale — applies to itself one more time: the record shape is safe to build against now, and the cost figures in it are not yet safe to reason about. That is a narrower hold than the one it replaces, and it is the honest one. Outcome: hold narrowed, still no ask for @danmtNothing is minted by this comment and no label moves. There is no question on this page for you, which is why it carries no return date and still needs none. The wake condition is re-stated so it cannot go dark again: this thread wakes when #585 lands an observed billed envelope, or when a fresh fleet corpus is collected — the last one was taken Triage, 2026-09-01. |
Uh oh!
There was an error while loading. Please reload this page.
Why this is a thread and not an issue
Its own three section bodies, quoted exactly:
An issue is a work order a builder must be able to execute without asking anyone anything (TRIAGE.md). No tasks, no acceptance criteria and five open questions is not a work order — and acceptance criteria are not optional furniture: they are the builder's definition of done and the reviewer's review spec, verbatim. Without them nobody downstream has a spec at all.
.ceremony/TRIAGE.mdnames this exact move in its "what you never do" list: "Mint an issue to 'discuss' something — that is a discussion." Triage did it anyway, five times, and this is one of the five.The
blockedlabel was the sharper problem. It means waiting on another issue or PR (LABELS.md) — which reads, to a scanning builder, as work that becomes claimable when #122 lands. It would not have been. This board wrote that lesson down itself, on #116: "the queue label is a dispatch signal, not a status annotation. A prose caution under areadylabel is not a caution, it is a trap." A prose caution underblockedis the same trap with a longer fuse.What survives, because it is good
The working context / state context split is the sharpest idea on this board about session cost, and it is a decision, not a question:
with the reasoning that
--resumecarries both undifferentiated, and a round-3 session that "remembers" round 1's head is the exact bug class the fleet spent 2026-07-28 killing (#114). That belongs in whatever eventually gets built, and it is why this thread exists rather than awontfix.The outcome: ask — and the questions are not answerable yet
Not by @danmt, not by triage, not by a builder. Every one of them is a question about a crossover point — does carrying notes beat a cold start, and after which round — and the engine currently reports no token or cost number at all. That is #122's job.
So this thread is held on #122, deliberately, for the reason the issue already gave: "writing the table now would encode a guess as a spec" (the same reason #92 sat blocked). When #122 lands and the numbers exist, triage returns here, asks the five questions against real figures, and mints from the answers.
Nothing is waiting on this thread. It gates no issue and blocks no builder.
The design record, as filed on #119 — verbatim
The waste
Every duty session is a cold start:
run_session()runsclaude … -p "$prompt"with no--continueand no--resume. A builder answering round 3 of a PR re-reads the repo, re-finds the same locations, and re-writes the same throwaway search scripts it wrote in round 1.That part is straightforwardly wasted. What is not obviously wasted is the re-reading of state — and the difference is the whole design.
Two kinds of context, and only one is safe to carry
mainlooks like now, which labels stand--resumecarries both, undifferentiated. A round-3 session that "remembers" round 1's head is the exact bug class this fleet spent 2026-07-28 killing: verdicts keyed to stale heads (#114), dependency declarations that never re-validate their own completeness (#96), a reviewer approving without re-reading.So the mechanism must carry working context and refuse to carry state.
Why a written working-memory artifact beats
--resumedanmt's framing — "each session writes work-related notes and a little map of what it has; instead of resuming you say continue doing X, here's some background" — is the better mechanism, on five counts:
--resumeis per-vendor and may not exist for three of the four — the same per-vendor wall the tier question hit.~/duty/work/already exists (install.sh:160) and is preserved across installs. A session transcript dies with a box rebuild or a gold restore.It is also the architecture this fleet already chose. #196/#91 ruled that the engine gathers context, the model transforms, and the engine keeps the artifact. This is that same shape applied to a different duty: the engine assembles working notes plus current state, the session works, the engine keeps the notes. One pattern, not two.
Not every duty wants this
Builders and reviewers: yes. Multi-round work on one deliverable is where working context is rebuilt over and over.
Triage: no — and the reason generalises. danmt's read is right: triage's answer is to make the issue carry enough context to be cheap to re-read. That is not a consolation prize, it is a better optimisation, because triage's output is every downstream session's input. An issue that names the file and line, states the class, and records what was already checked is paid for once by triage instead of once per builder per round.
There is evidence for this already. A builder picking up #114 does not need to locate
shared/lib/duty-review.sh:116, re-derive that the bug is a class rather than a site, or check whether the file the reporter named still exists onmain— triage did that once. The issues written on 2026-07-28 are expensive to produce and cheap to consume, and that is the trade this issue is about, made at a different layer.Open questions — genuinely open
~/duty/work/<repo>#<num>/is the obvious home, but a worktree-local file has better locality and worse durability.Tasks
Deliberately none yet. This issue is the design; the task list is written once #122 says what the sessions actually cost.
Acceptance criteria
Deliberately none yet — same reason. What it must eventually satisfy:
Dependencies
Blocked by #122. That issue's
--output-format jsonchange returns both the usage numbers that would justify this and the session id that a--resumevariant would need — one change, two unlocks. Neither this nor #121 gates0.1.0.Related: #120 (token-efficiency techniques), which is the same class of speculative optimisation blocked on the same numbers.
All reactions