Techniques for token efficiency — a candidate list, deliberately unranked until the numbers exist #233
Replies: 9 comments
Update — one of the six candidates stopped being a guess today, and its evidence expires at tomorrow's rolloutThis thread's outcome line says the list is held on #122 and closes:
That was true when it was written yesterday and it is now true of five candidates, not six. Re-read from the board rather than from this thread. Candidate 6 was never a token question
Candidates 1–5 all trade an unknown saving against a correctness risk, and none of them can be ranked without a per-session token figure — which is exactly what #122 does not yet emit and why this thread holds. Candidate 6 is a different unit. It does not ask what a session costs; it asks how many sessions should not have run. That is countable from #290 is the vehicle, it is
Two things made that possible and neither existed when this thread was written. #256 landed — a zero-action session ending in a question no longer reports Why this is worth a comment rather than a shrugThe evidence is perishable and the deadline is physical. #290's phase 1 exports each box's And the ranking act this thread reserved is the one #290 performs. This thread's own words: "That first act is triage's, not a builder's: reading measurements and ranking candidates is what produces the work order." #290's deliverable is a written report whose findings are each filed as issues, plus a proposed What does not moveCandidates 1–5 stay held on #122, unchanged and for the unchanged reason. #290's corpus carries counts, durations and outcomes — no token or cost figure anywhere in it — so terse prompting, tool-result discipline, a code graph, context compression and cache alignment remain unrankable, and committing to one would still encode a guess as a spec. #122 is open and The ordering principle stands, and it is what makes this split legitimate rather than opportunistic: ranked by how cheaply each can be tested, not by expected saving. Candidate 6 just became the cheapest to test by a wide margin, which is where the list said it would earn its place. Outcome, restated: ask — held on #122 for five of six candidates, and no longer held for the sixth. Nothing is waiting on this thread, and it gates no issue. Triage, 2026-08-02. |
Candidate 6's vehicle is claimed and in flight — the deadline this comment warned about is being metShort update, from label events rather than from this thread. #290 is no longer The split holds exactly as stated. Candidates 1–5 stay held on #122 — the report's corpus carries counts, durations and outcomes and no token or cost figure at all, so nothing in it can rank terse prompting against tool-result discipline. Candidate 6 is the one that stopped being a guess, and its measurement is what #290's phase 2 performs. One caveat worth having on the record before the numbers arrive, because it is the kind of thing a reader infers wrongly from a good report: a low zero-action rate would not retire candidate 6. The rate counts sessions that ran and did nothing; the candidate's thesis is about sessions that should never have been spawned, and the #286 loop is the proof that those two are different — six reviewer sessions at an unchanged head, each of which did something, all six of which should not have existed. Triage, 2026-08-02. |
The corpus landed, candidate 6 has a measured instance — and it is five times smaller than the thing beside itMy comment above said candidate 6 stopped being a guess and that its measurement is what #290's phase 2 performs. Phase 2 has not run — #290 is still The hold on candidates 1–5 stands, and is now a measurement rather than a claimThere is no token or cost figure anywhere in the corpus. So nothing in it can rank terse prompting against tool-result discipline, and #122 ( Candidate 6 — "not spawning the session" — priced in wall-clockTwo things the corpus can price, both from The #311 flood. 59 And the thing next to it, which is bigger. Nine sessions in the same window ended
Candidate 6's own logic points at the second table, not the first. A session that shouldn't have been spawned is the cheapest possible saving — and a session that ran to its timeout and produced nothing is the same saving, five times over, in the same two days. The flood is the more offensive shape (93 comments, a legible loop) and the smaller number. The correction this makes to my own comment aboveI wrote: "a low zero-action rate would not retire candidate 6 … the candidate's thesis is about sessions that should never have been spawned." That still holds. What I did not anticipate is the direction of the surprise: the measured zero-yield burn is dominated not by sessions that should never have been spawned, but by sessions that were correctly spawned and then did not finish. That is a different defect class with a different fix, and it is not on this list at all. Not filed, and the reason is the one this thread already established. Findings from this run route through #290's finding set, and a triage mint that front-runs the report inverts what that issue exists to do. #314 is the exception rather than the precedent — the operator pulled it onto the cut. The timeout finding is recorded here and on #290 so the report does not re-derive it. Outcome unchanged: ask — held on #122Wall-clock is not spend, and this comment ranks nothing. When per-session cost exists, triage ranks the candidates here against real figures and mints the winner — as this thread has said since 2026-07-28. Nothing is waiting on this thread. Triage, 2026-08-03. |
Candidate 6's vehicle landed and closed — the figure I priced it at is superseded by a wider boundary, and the hold on 1–5 is now a merged statement rather than triage's assertion#290's report merged at The hold on candidates 1–5 stands, and stopped being something triage asserts
That is the merged report's own sentence. #122, re-read from its label events rather than from this thread: Candidate 6's number is superseded — by a boundary, not by a correctionThis morning I priced it at 59 resume sessions / 61.7 minutes, measured from the last commit at
on a boundary that starts at the And do not carry my "89% of all resume duty" percentage onto it. The report's fleet-wide resume total is The direction-of-surprise correction I made this morning is now the report's own ranked findingI wrote that the measured zero-yield burn is dominated not by sessions that should never have been spawned but by sessions that were correctly spawned and did not finish. The report puts a multiplier on it:
Both boundaries stated in one sentence, which is what makes the comparison legitimate. That class is not on this thread's candidate list and is not a token question either — it is routed to the sentinel's layer-one indicators, recorded on #236. Two pointers in my comment above are stale
Outcome unchanged: ask — held on #122 for candidates 1–5Wall-clock is still not spend, and nothing here ranks a technique. Candidate 6 has been measured, its instance has a filed and claimed fix, and the class next to it turned out larger — none of which prices a single one of the other five. When per-session cost exists, triage ranks the candidates here against real figures and mints the winner, as this thread has said since 2026-07-28. Nothing is waiting on this thread, and it gates no issue. Triage, 2026-08-03. |
Candidate 6's fix shipped, and the issue candidates 1–5 are held on was closed — one of those is delivery and the other is bookkeepingTwo board events since my comment at 1. Candidate 6 — not spawning the session — is the first item on this list to ship#314 closed at
So the one candidate this thread could price is also the one with a landed fix, thirty-six hours after it stopped being a guess. And its generalisation now has a window: #329 ( 2. Candidates 1–5: the issue they are held on closed, and nothing was delivered by that#122 closed at The hold's reason is exactly what it was. There is still no token or cost figure: none in the engine, none in the corpus, and the merged report's own sentence stands — "the corpus records no token or monetary cost per tick." Terse prompting, tool-result discipline, a code graph, context compression and cache alignment remain unrankable, and committing to one would still encode a guess as a spec. What moved is the address. The wake is no longer "when #122 lands"; it is #327 ( 3. The class this thread found and does not own now has a carrier tooMy comment this morning recorded the direction-of-surprise: the measured zero-yield burn is dominated by sessions that were correctly spawned and did not finish — nine It is now on #327's to-mint list, as the "tick-health surface: last-tick age, skip/hold counts, outcome streaks — the numbers the 2026-08-01/02 report mined by hand", and on #236 as indicator 1. Recorded here so this thread's list is not the place someone goes looking for it. Outcome unchanged: ask — held on the numbers for candidates 1–5The ordering principle is untouched — ranked by how cheaply each can be tested, not by expected saving — and nothing here ranks a technique. When per-session cost exists, triage ranks the candidates against real figures and mints the winner, as this thread has said since 2026-07-28. The place that produces the figure is now #327's window. Nothing is waiting on this thread. Re-derived at this write with Triage, 2026-08-03. |
Two candidates were added to the fleet's efficiency backlog tonight — on an epic, not on the list that claims to be the listThis thread's standing promise is that it is the candidate list, and that triage ranks the candidates here against real figures and mints the winner when per-session cost exists. Two new ones landed at Candidate 7 — risk-tiered panelsEvery PR costs Candidate 8 — per-duty model-tier routingNamed in the same fold-in as a sibling lever, "recorded in #121's notes and on no epic's surface." Measured: it has a full design record of its own at #92 — the Where they sit against this list's ordering principleThe principle is ranked by how cheaply each can be tested, not by expected saving — and candidate 7 is the first arrival that genuinely tests cheaply, because its premise is a board measurement rather than a token one. Verdict counts and finding counts are public and re-derivable by anyone; per-session cost is still the thing that does not exist. So triage re-derived its premise instead of repeating it, from the three PRs the fold-in cites, and it did not survive intact — the full measurement is on #329, where the mint will read it. In short: #342 is the clean instance (docs-only, four verdicts, nothing found); #315 is not docs-only by path — it changes That is this thread's own constraint arriving on the newest candidate, and it is why the sentence at the top of this list still governs:
A docs tier is a deliberately lossy optimisation on the detection axis. It may well be worth it — idle fourth verdicts on Candidate 8 is unmoved and unrankable, for this list's unchanged reason: #92 states its own gate correctly — without measured per-kind burn the table encodes a guess. Same gate as candidates 1–5, at the same address. Outcome unchanged: ask — held on the numbersCandidates 1–5 and 8 are held; candidate 6 shipped in Nothing is waiting on this thread. Re-derived at this write by running Triage, 2026-08-03. |
Candidate 6 has a second measured instance, and it shipped this morning — the only line on this list that has ever produced workOne delta since my comment of 2026-08-03 1. Sixteen sessions that could not succeed, in thirty-five minutes#388 was filed 2026-08-06 from a live incident on
That is candidate 6 exactly — a session that shouldn't have been spawned is the cheapest possible saving — at sixteen instances of it, unbounded, with two PRs unreviewable throughout. 2. What shipped, and what it is worth on the incident's own shape
On the shape #388 recorded, that is three doomed sessions instead of sixteen, the remainder replaced by one probe per tick, and a The second blindness the issue names is closed too, and it is the enabling half rather than a footnote: 3. What this does to the ordering principle: nothing, and that is the findingBoth of candidate 6's instances arrived the same way and neither came from this list. #91 was found by noticing what handoff was spending; #388 was found by @danmt reading raw session logs on a box, which the issue's own first line calls the finding. The list's ordering principle — rank by how cheaply each can be tested — still has not been exercised once, because the thing it ranks against still does not exist: the engine reports no token or cost number, exactly as the design record said on 2026-07-28. So the honest status of this page is two-tiered, and worth stating plainly rather than leaving implicit:
Candidate 6's own sentence is the one to carry forward: "Look for the next one of those before optimising the ones that should exist." It has now paid twice, and both times the search was an operator or a session reading a log rather than a technique being ranked. 4. The address of the hold is nearer than it has ever been
Nothing here ranks a technique, and no candidate is minted. An issue enters a release by decision, not by triage's default. Outcome unchanged: ask — held on the numbersNothing is waiting on this thread, re-derived at this write by running the pinned Triage, 2026-08-07. |
The number this list has been unranked-until since it opened now has a decided shape — and a caveat that bounds how the ranking may be readOne delta since my comment of 2026-08-07 1. What was ruledThe reading taxonomy,
So the figure this thread waits for is not a number, it is tokens by model per session, priced at read. For a ranking that is strictly better than a stored cost: candidates can be re-scored under a corrected price table without re-capturing anything. 2. It settles candidate 8 outright — worth saying, because I filed it as unmeasurableCandidate 8 — per-duty model-tier routing was added here on 2026-08-03 with a design record of its own at #92 and no way to evaluate it. Candidate 7 — risk-tiered panels is priced off the same rows — 3. The caveat, which bounds the ranking rather than delaying itRuled by @danmt in the same session: credential identity on the session row is an operator-declared label, defaulting to the box's own identity. The alternative — a derived non-reversible in-box digest, or the vendor's own account identifier — was considered and declined as overbuilt for an interim, with the residual risk accepted knowingly and in writing:
What that costs a ranking run from this thread: the unit is the declared label, never the box. Two boxes on one credential with distinct labels split one spend across two rows, and no assertion anywhere can catch it — the ruling says so rather than leaving it to be discovered. A ranking that reads per-box totals as per-credential truth under-prices precisely the candidates that touch the busiest shared identity, which on this fleet is the reviewer panel, and therefore candidate 7 — the one candidate on this list whose whole claim is about panel size. So when the ranking runs it states the unit it used, and the declared labels are checked against reality before it runs rather than after. 4. Where this leaves the threadAnswered; still held, and the hold's address is a schema now rather than a window. The ranking is still due when #327 ships capture — that window is still gated on #346's close and Triage 2026-08-09 |
The hold has lifted — the grep this thread has been unranked-until now returns hits, and the ranking is dueTriage 2026-09-01, re-measured at This thread carries no return date, correctly — its wake is an event: "the ranking is still due when #327 ships capture." That event has fired, and nothing woke this thread when it did. It is 8 days late for a structural reason worth naming: an event-keyed wake whose event nobody is watching reads as "nothing due" forever, on every surface — no sweep, no nudge, and not the Capture shipped. The measurement, run exactly as my comment of 2026-08-09's sibling ran itThe
Carried by The issues that did it are all closed: #475 (capture — #122's re-mint), #485 (vendor quota), #483 (the vitals probe), #538 ( So the condition this thread set for itself is met: the ranking is no longer waiting on a schema, on a window, or on a decision. It is waiting on numbers, which is a different and much shorter wait. The one thing still genuinely missing, and it is not a schemaThe tree can now emit a cost figure; nobody has observed a billed one. #585 — open, That bounds how this thread's ranking may be read, and it does not block writing it. The candidate set is still all eight and the ranking still runs over all eight or it is not the ranking this thread promised — but the ranking method, which the 2026-08-09 comment already said could be written before the numbers land, can now be written against a real emitted record shape rather than a specified one. The figures that fill it in are #585's, and the fleet corpus that would price them at scale is a collection nobody has re-taken since Outcome: hold released, still no ask for @danmtNothing here is minted by this comment and no label moves. There is no question on this page for you, which is why it carries no return date and still needs none. The wake condition is re-stated so it cannot go dark again, since the last one did: this thread wakes when #585 lands an observed envelope, or when a fresh fleet corpus is collected — whichever is first. Both are events with a board entry behind them now, which the previous wake condition did not have. Triage, 2026-09-01. |
Uh oh!
There was an error while loading. Please reload this page.
Why this is a thread and not an issue
Its own header, quoted exactly:
And its title says the rest: "a candidate list, deliberately unranked". An unranked list of guesses, explicitly marked do-not-build, is a discussion — TRIAGE.md's "what you never do" names the move directly: "Mint an issue to 'discuss' something — that is a discussion." Triage did it anyway.
The
blockedlabel made it worse rather than safer.blockedmeans waiting on another issue or PR (LABELS.md), so on a board scan this read as work that becomes claimable the moment #122 lands. It would not have. #116 records the lesson in this repo's own words: "the queue label is a dispatch signal, not a status annotation. A prose caution under areadylabel is not a caution, it is a trap."What survives, because it is good
The ordering principle — ranked by how cheaply each can be tested, not by expected saving — and the constraint that frames all five candidates:
That sentence should govern whatever is eventually picked, and candidate 2 (tool-result discipline — "probably the largest single line item and the least glamorous") is very likely the right first move on the numbers.
The outcome: ask — held on the numbers, and the first move is triage's
The question this thread owes an answer to is not "which technique?" It is, in the issue's own words:
That first act is triage's, not a builder's: reading measurements and ranking candidates is what produces the work order, and it cannot be delegated to whoever picks the ticket up. So this thread is held on #122. When per-session cost is visible, triage ranks the candidates here against real figures and mints the one that wins.
Recorded honestly: danmt's starting position is "I have a gut feeling we're wasting an absurd amount of tokens." Probably right. Still not measurable, and that has not changed since 2026-07-28.
Nothing is waiting on this thread.
The design record, as filed on #120 — verbatim
Why a list and not a plan
Each technique below trades something. Without #121's numbers there is no way to tell which is buying a real saving and which is buying complexity at the cost of correctness — and this fleet's failure mode is not "too slow", it is "silently wrong", so a lossy optimisation in the wrong place is worse than the spend.
Ordered by how cheaply they can be tested, not by expected saving.
Candidates
1. Terse prompting. Strip prose from the rendered prompts; keep the contract. Cheapest to try, easiest to measure, and the only one with no correctness risk beyond a prompt that is now ambiguous.
shared/prompts/*.txtare versioned templates, so an A/B across boxes is mechanical once #122 reports per-session cost.2. Tool-result discipline. Much of a session's spend is likely tool output the model never uses — full file reads where a range would do, unbounded
grep, whole diffs. Not a technique so much as an audit; probably the largest single line item and the least glamorous.3. A code graph or index. Pre-build a map of the repo so sessions stop re-deriving where things live. Real saving; two real costs — it must be maintained, and a stale index is worse than none, because the agent trusts it. Interacts directly with #119: a shared index is a fleet-wide version of that issue's per-deliverable working notes.
4. Context compression. Summarise long inputs before they reach the model. Highest ceiling, highest risk: compression is lossy by definition and the loss lands where nobody looks. Would need the same treatment #119 gives state context — never compress the volatile half.
5. Prompt-cache alignment. Order the prompt so the stable part is a stable prefix. Costs nothing at runtime and may already be partly true by accident; worth knowing before changing prompt structure for other reasons, since #119's notes-injection could defeat it.
6. Not spawning the session. #91's thesis, and it belongs on this list because it beat every technique here: a session that shouldn't have been spawned is the cheapest possible saving. Handoff was spending a full agent session and a repository clone to make two API calls and write one paragraph. Look for the next one of those before optimising the ones that should exist.
What is already working, unglamorously
Dense triage issues. An issue that names the file and line, states the class, and records what was already checked is paid for once by triage instead of once per builder per round. A builder on #114 does not re-locate
duty-review.sh:116, re-derive that the bug is a class, or re-check whether the reporter's named file exists onmain.Worth stating because it is the counterexample to the whole list: the largest token saving found so far came from writing things down properly, not from a technique.
Dependencies
Blocked by #122. Nothing here should be built, and no candidate ranked, until per-session cost is visible. The first task after #122 is not to pick one — it is to look at the numbers and find out whether the gut feeling is right and where.
Related: #119 (working memory vs state memory), which is the same class of optimisation at the session layer and blocked on the same numbers.
All reactions