docs(s-h): P14 /context addendum — the operator paste that answers DECISION-NEEDED #3 - #1249
Merged
Conversation
added 13 commits
August 7, 2026 03:22
…EDED #3, and falsifies the 4 B/token convention The operator ran `/context` post-merge and supplied the output, taking Option A of DECISION-NEEDED #3 (§0a). Recorded as a new §8 addendum rather than an in-place rewrite: the measurement history must read "unknown at stage close -> known 2026-08-07", not as though the split had been available all along. §0a is kept verbatim, annotated ANSWERED. Two findings, in order of consequence: §8.1 — the seed's binding 4 B ~ 1 token conversion is FALSIFIED. Seven files carry both a `wc -c` byte count and a harness-reported token count; aggregate 77,156 B / 29,464 tok = 2.62 B/token. Every 4 B/t figure in this patch and its sibling is low by ~1.53x, and because row 5 (harness remainder) is computed BY DIFFERENCE, the remainder is correspondingly HIGH — a first-order restatement puts it near 52%, not 68.4%. Figures are left as published and the correction is recorded as owed, not made: the row-1 file set is the pre-S-G resident set while the ratio was measured on the current one, so they are not the same population, and re-deriving §2 on a new constant is a re-measurement beyond an addendum. Raised as DECISION-NEEDED #4. §8.2 — the reported percentages sum to 105.6% because the two `(deferred)` rows are counted but NOT resident. The identity confirms it exactly: 334.6k - 276.4k = 58.2k resident, and the non-deferred rows sum to 58.2k. Half the resident head is memory files (29.4k of 58.2k), of which two documents carry a third of everything (repo CLAUDE.md 9.3k + ai-laziness-traps.md 9.8k). ToolSearch deferral withholds 58.1k — almost exactly what the entire resident head costs, which is the number §3's "preserve what already works" lacked. §8.3 — closes TWO of the four `UNMEASURED — channel absent` rows, not four: 5c (MCP tool schemas, 8.4k) and 5e (skills 8.9k + custom-agent listing 1k). 5d stays open (`/context` does not itemise server instructions apart from tool schemas) and row 9 stays open (a different population: "Custom agents" counts registered agent types, not the repo's agents/ directory). Neither was filled from the nearest plausible neighbour — that is T-SH-A working, not a shortfall. Revised partition 14 / 11 / 2 / 1, counted from the table. All count-claims re-swept by class after the edit rather than site-by-site (the W-9 lesson from the round-3 fidelity audit): table recount gives 14 rows and exactly 2 carrying the literal marker; the stage-close claims of "four" are retained as historical and each carries its revision inline. Coverage: n=1, an orchestrator seat in a worktree with five rule files injected; a fresh main-checkout or subagent seat has a different resident set. All figures are the harness's own estimates at its own rounding; no tokenizer was run. Prior-art: skipped — post-merge measurement addendum to an existing research patch, no new capability
MAJOR — §8.2 claimed `ToolSearch` deferral "roughly doubles the usable budget". The snapshot cannot support that: window 1m, free space 665.4k, so making the 58.1k deferred schemas resident moves free space to ~607.3k (-8.7%). What doubles is the resident HEAD (58.2k -> 116.3k). Restated to the measure the snapshot actually bounds; the supported neighbouring claims (58.1k is about the size of the whole head; still the most expensive available regression) are kept. MINORs, all from the same cold seat: - §0a heading was present-tense "five blocks stay unpriced", false after the update -> marked "(as at stage close) … stayed", with the current count (three: 5d, 9, row 8's injected form) stated in the ANSWERED block and again in §7. - "five rule files" contradicted the patch's own table -> four, with the four named and the other three memory files identified. - DECISION-NEEDED #4 Option A pointed at the sibling's "§5/§9"; the sibling has no §9 (it runs §0-§8) -> corrected to its actual 4 B/t sites, §5, §7 and §8. - The 5d basis asserted server instructions "sit inside the system-prompt region"; the capture establishes only that /context does not itemise them apart from tool schemas -> the locational claim is dropped, since asserting a region is the estimate T-SH-A forbids. - rows 1-4 restatement read 30,163, which reproduces from neither derivation route -> 19,719 × (4/2.6187) = 30,120, remainder 62,340 - 30,120 = 32,220, share 51.7%. All count-claims re-swept by class after the edit: table holds 14 rows with exactly 2 carrying the literal UNMEASURED marker; every surviving "five" is either historical-and-marked or refers to the item-4 probe's five files, a different subject. Prior-art: skipped — review-absorption edit on an existing research patch, no new capability
…check, surface the channel disagreement Round 2 confirmed all six round-1 findings closed and re-derived every §8 figure independently, then found two MAJORs the addendum had not noticed about its own effect on the rest of the file. MAJOR 1 — the §1.7 backward-check asserted SWEPT-CLEAN using figures this same commit restates. Both verdicts re-adjudicated in place rather than left standing: - ADR-3: the "inside ADR-3's stated band" clause was wrong when written — the band is 29-39% and both measurements (27.8% / ~21%) fall BELOW it; under §8.1's conversion the same share moves to ~47%, outside on the high side. Now GAP-FOUND, direction unresolved pending #4. - the spec's P14 row: "the row's arithmetic holds" is true only under the 4 B/t constant it was computed with, since §8.1 restates the same seat at 51.7%. Now HOLDS-CONDITIONALLY on #4B. MAJOR 2 — one seat, two irreconcilable harness figures, previously unflagged. By difference the main seat's remainder is 69,300 of 89,019; /context's categories matching row 5's own definition sum to 28.8k for that SAME session, and neither 28.8k nor 86.9k (adding deferred schemas back) reaches 69,300. The totals disagree the same way: 58.2k resident vs 89,019 first-turn billed, gap ~30.8k. New §8.5 states the disagreement, offers the dispatch-prompt hypothesis explicitly as unmeasured (§0 defines the channel as "resident head PLUS its dispatch prompt", and rows 1-4 never subtract it; this session opened with /orchestrator, which injects a whole SKILL.md body), and draws the consequence that matters: by-difference systematically OVERSTATES the remainder, because anything it cannot attribute to rows 1-4 lands in row 5 by construction. Raised as DECISION-NEEDED #5 with three options including "measure the gap directly". Not resolved here. MINORs: - §4 was the only section the revision sweep had skipped. R1 now carries a PERFORMED block (the paste happened; two of four rows closed, not four; S-D′ no longer has to park). R5's conclusion is REVERSED with its reasoning shown — its "next lever is harness-side" is contradicted by memory files being 50.5% of the resident head and repo-owned. - rows 5c/5e now carry the seat annotation: orchestrator MAIN seat, n=1, not the 62,340-tok subagent seat the table is sized against, with an explicit do-not-sum-against-row-5. - The headline now warns that both its percentages are contested, naming #4 and #5. Count-claims re-swept: 14 table rows, exactly 2 carrying the literal UNMEASURED marker. Prior-art: skipped — review-absorption edit on an existing research patch, no new capability
… the 600-line gate Round 3 confirmed round-2's MAJOR #2 (channel disagreement) and MINOR #4 (seat annotations) fully discharged, and re-derived every §8 figure independently. It then caught the replacement figures themselves. MAJOR — the ADR-3 re-verdict swapped one unsupported number for another: `29,464 / 62,340 ≈ 47%` divides a MAIN-seat /context numerator by the SUBAGENT-seat by-difference denominator — exactly the cross-seat, cross-channel mix this same commit forbids at rows 5c/5e and that §8.5 declares irreconcilable. Restated within one channel: 17,363 × (4/2.6187) = 26,522 = 42.5% of the 62,340-tok seat, or ~26.5% against ADR-3's own ~100k denominator. Both readings put the repo-owned share BELOW the 29-39% band, not above it, so the verdict is now "GAP-FOUND — measured low, consistently across the conversion change" instead of "direction unresolved". MINORs: - §8.1 gave "two reasons, both binding" for not reconverting §2; one was FALSE. Concatenating the five files the ratio was measured on gives 69,453 B and row 1's published 17,363 est-tok is exactly 69,452 B / 4 — the same population, byte for byte. The claim is withdrawn in place and the surviving reason (re-derivation is beyond an addendum) is named as the only one. An unverified escape clause is a stronger shield than the correction it blocks, and this one was steering DECISION-NEEDED #4. - DECISION-NEEDED #5 Option A's "wrong by roughly 2.4x" over-extended: 2.41x is the main-seat ABSOLUTE; the share moves 77.8% -> 49.5%, i.e. 1.57x, and the subagent-seat 68.4% is untouched because /context cannot run inside a subagent. - §8.2 called the whole 29.4k memory block repo-owned; 2,764 of it is host-side (~/.claude CLAUDE.md 964 + MEMORY.md 1,800 = §2 rows 2 and 3). Repo-owned is 26,700 = 45.9% of the head. - The §4 sweep had reached R1 and R5 but not R4, whose premise the paste contradicts: R4 rests on the harness truncating the skills listing "to a ~2k budget", while /context measures the injected block at 8.9k — essentially the un-truncated source-side ~9.1k. Surfaced for S-I, not re-derived here. R4's "129 SKILL.md files" also carries no reproducing command and a recount gives 112, so the population is marked UNVERIFIED. Structural: absorbing the above pushed the patch to 602 lines, over the repo's 600-line markdown gate. Trimming to 599 would be gaming the gate, so §8 is split into a companion patch, 2026-08-07-s-h-p14-context-addendum.md, with §8.x numbering preserved so every cross-reference already written stays valid. Parent 435 lines, addendum 196. Prior-art: skipped — review-absorption edit plus a size-driven split of an existing research patch, no new capability
…e 13) The pre-push principle-13 gate correctly rejected the new patch: a research patch must carry an actual §1.7 self-review, not merely name the section. Added Forward + Backward + T15. The backward-check is a real outward sweep, not a restatement of this diff — the change class is "a post-merge artefact that revises figures already published in a merged research patch", and six surfaces are verdicted, of which four are GAP-FOUND and left unedited by ownership: - the sibling p3d-p11 patch shares the falsified 4 B/t constant at its §5/§7/§8, so §8.1 applies to it identically — named in DECISION-NEEDED #4's Option A as required scope; - the token-economy spec's tag convention (the constant under one of its tags is wrong); - ADR-3 (repo-owned share measures below its 29-39% band under BOTH conversions); - the S-A profile patch (closed historical artefact, its authoring session owns it). T15 records the reflexive fact that this file exists only because the parent hit the 600-line markdown gate — a document about document cost split by a size discipline. Prior-art: skipped — self-review section required by principle 13 on an existing patch, no new capability
… round's own replacement figures The split is sound (parent 463, addendum 261, all 43 §8.x cross-references resolve) and the addendum's §1.7 backward-check verified as a real outward sweep. But three of round 3's four replacement figures were themselves defective, plus a new challenge block that reversed a downstream premise on an invalid comparison. MAJOR — "Both readings put the repo-owned share BELOW the 29-39% band" is arithmetically false: 42.5% > 39%. And 42.5% is a share of the 62,340-tok SUBAGENT seat while ADR-3's band is denominated on ~100k, so it is not band-comparable at all. Round 3 replaced a cross-SEAT mix with a cross-DENOMINATOR one. Now stated from the directly measured figure with both traps recorded inline so it is not re-derived wrongly a third time. MAJOR — 26,522 was derived by applying §8.1's SEVEN-file aggregate ratio (2.6187, inflated by the one host-side Russian-text outlier at 3.32 B/t) to row 1's FIVE-file population, while the addendum measures that exact population directly at 26,700 (five-file ratio 2.6012). One commit, two values for one block. The measured figure now supersedes the derivation: 26,700 = 26.7% against ~100k (band-comparable, below the band) and 42.8% of the subagent seat (not comparable). MAJOR — round 3's "29.4k is not repo-owned" fix was applied in §8.2 but not swept: §8.4 (the S-D′-facing ranking section) and §4 R5's REVERSED note both still read "29.4k, 50.5% repo-owned", overstating the own-able block by 2,764 tok at the one site a downstream stage reads. Both fixed to 26,700 = 45.9%. Third site of the same class: "six ASCII-dominant repo files" counted host-side MEMORY.md as a repo file. MAJOR — the R4 CHALLENGED block concluded the skills listing "appears not to be truncated at all", comparing the /context-measured 8.9k against a ~9.1k figure that is a 4 B/t estimate this same commit declares low by 1.53x. In one constant: 41,057 B / 2.6187 = 15,678 tok, so 8.9k is ~57% of source; independently the snapshot lists 74 entries against a 112-file population, ~66%. Both channels say REDUCED. The supported half survives — the ~2k budget premise is wrong by ~4x — and that, not "no truncation", is what is routed to S-I. MINORs: §4 R2's "until then / which R1 would settle" was stale once R1 discharged (now PARTLY SETTLED, with the evidence stated as non-conclusive and the row keeping its UNMEASURED pricing rather than gaining a "0"); §7's S-I-kickoff backward-check verdict was not re-adjudicated although this commit moves that kickoff's premise (now GAP-FOUND, routed not edited); the T3 demand for a reproducing command was applied to the 129 being corrected but not to the 112 correcting it (command now published beside it). The addendum's §1.7 now records the method failure rather than only the rows: four rounds, four sweeps driven by the last review's list, each re-failing on whatever the list omitted — T21 in its own-work form. Prior-art: skipped — review-absorption edit on existing research patches, no new capability
…them a fourth time Round 5 found 2 MAJOR, both again in the previous round's replacement figures. That is four consecutive rounds where a hand-revised quantitative claim was itself defective, so this round changes method: the unsupportable claims are withdrawn rather than corrected again. MAJOR — the ADR-3 verdict was denominator-SELECTED, not measured. 26,700 has four defensible denominators and they disagree in direction: 29.99% of this seat's own 89,019 first-turn total (INSIDE the 29-39% band), 26.6% of the 60-session median (below), 45.9% of the /context resident head (above), 42.8% of the subagent seat (above). Rounds 3-5 each picked one and each pick was defective — cross-seat, then cross-denominator, then ratio-transferred-across-populations. The verdict is now WITHDRAWN with all four denominators tabled and no verdict issued, and the choice raised as DECISION-NEEDED #6 (which cannot be settled independently of #5, since the options differ precisely by the ~30.8k dispatch-prompt gap #5 records). MAJOR — the "74 listed entries / 112 files = 66%" corroborating channel is WITHDRAWN entirely. The numerator is provably not a subset of the denominator: the two largest listed entries in the capture, dataviz (~380) and claude-api (~360), have no SKILL.md anywhere, as do >=14 other built-ins. The denominator is an unfiltered find carrying marketplace/cache duplicates, vendored node_modules files, worktree copies, packages/core fixtures and uninstalled catalogue rows. A ratio across two different sets measures nothing; publishing it would be the estimate-dressed- as-measurement T-SH-A forbids. MINORs: the "~57% of source" precision is withdrawn to direction-only — it swings 56% to 87% across the four conversion constants in play, and the SKILL.md corpus is itself multi-byte-heavy (six skills carry Russian descriptions), so no constant is defensible for it without measuring that corpus. The 112 recount is no longer offered as a correction: publishing the command is necessary but not sufficient, since the command must already exclude what the claim is not about. measure-always-on.sh's "21-28%" gained the re-adjudication marker every sibling surface had. The §1.7 note previously NAMED T21 while committing it. It now states plainly that this round's sweep was list-driven too, that its hunks map one-to-one onto round 5's findings, and that the class-driven counter T21 prescribes is what the five cold audit rounds have been doing while the author-side sweep never became class-driven. It also records the second method finding: three attempts to repair one comparison failed because the comparison had four denominators, and the correct response was withdrawal. Class sweep applied to the withdrawal itself: every site carrying a listing share was found by grep and corrected, not only the one the audit named — the §7 S-I re-adjudication repeated the withdrawn 57%/66% pair and now reads direction-only. Prior-art: skipped — review-absorption edit on existing research patches, no new capability
…ate ranges, never magnitude words Round 6 confirmed the round-5 withdrawal is complete (57%/66% survive nowhere; 74/112 only inside their own WITHDRAWN notices; no third site) and re-derived all four tabled shares as correct. It then found three MAJORs, all again in this round's own replacement wording. MAJOR — withdrawing "~57% of source" to "a minority of source" INVERTED the claim. Under every constant the same note lists, the injected 8.9k is 56.4% / 56.8% / 72.0% / 86.7% of source — a majority — and 51.4% against the pre-S-I byte count. A magnitude word is not a weaker form of a number, it is a different claim. Both sites now carry the explicit range and NO magnitude word; the withdrawal rule is stated so the next editor does not substitute another adjective. MAJOR — the new measure-always-on.sh re-adjudication claimed the measured 26,700 supersedes the "21-28%" pair. Wrong on the NUMERATOR, not the denominator: 26,700 is the pre-S-G five-file set (pinned byte-for-byte in §8.1) while the "~21%" member is the post-S-G set. No denominator choice repairs a numerator mismatch, so no restatement is offered at all — the bound is unverified here and both the surface and the post-S-G measurement stay S-E's. MAJOR — "Options A/B and C differ precisely by the ~30.8k gap" holds only for A (89,019 − 58,200 = 30,819). B differs by 42,329 and is a 60-session median set against a gap measured on one session, so B compounds #5 with a population change rather than restating it. Corrected in place. MINORs: the "six skills carry Russian descriptions" clause is DROPPED rather than corrected — two greps disagreed (6 vs a repo count polluted by node_modules), and the sentence two lines above faults another figure for lacking a reproducing command, so publishing an unverifiable one there was the same defect. "#6 below" pointed above. DECISION-NEEDED #6 is now propagated to every enumeration that had stopped at #5: the §2 headline warning, R5's REVERSED note (which quotes 45.9% — one of #6's four tabled options, now labelled as such), the §8 pointer, the addendum header and its §1.7 obligation count. The §1.7 note also records that this round's two records disagreed about whether the sweep found an unnamed site: the commit message was right, the paragraph was wrong. The S-I re-adjudication was found by the author's own class grep. Honest summary now stated: list-driven for five rounds, class-driven for exactly one item. Prior-art: skipped — review-absorption edit on existing research patches, no new capability
…ad of restating it an eighth time Seven cold rounds, seven REVISEs, and rounds 4-7 each found the MAJOR in the PREVIOUS round's own replacement wording. The class never changed: deriving a quantity across mismatched populations, denominators or conversion constants. This round stops deriving rather than deriving better. MAJOR (round 7) — the injected-vs-source share published last round has exactly the defect the adjacent paragraph withdraws another channel for: its numerator is the harness total for 74 LISTED entries, its denominator a byte sum over 129 SKILL.md FILES, and the same note proves those populations differ (dataviz ~380 and claude-api ~360 are in the numerator and have no SKILL.md at all). All four attempts at that share — a 66% population ratio, a ~57%, a "minority", a 56.4-86.7% range — are now withdrawn together with one stated reason, and NO figure, range or magnitude word replaces them. What survives is the comparison that needs no denominator: 8.9k measured against the spec's ~2k budget. MAJOR (round 7) — this file established last round that 26,700 is the PRE-S-G five-file set, then used that fact in exactly one place while three others kept building on it: §8.2 reading 1 and §8.4 ranked S-D′'s "top-ranked own-able target" from it, and R5 REVERSED overturned a recommendation with it. The dominant file in that measurement, ai-laziness-traps.md (26,387 B), is one S-G already replaced with a digest a quarter its size. So no ranking, no "top lever" and no share of the current head is derived from the snapshot; what it establishes is the narrower and still-useful claim that repo-owned memory files were the largest resident CLASS on the seat measured. R5's reversal is downgraded to SUSPENDED: its premise moved, but its opposite is not established either, and a post-S-G re-measurement is S-E's. MAJOR (round 7) — §8.4 still said "#4 and #5", the site round 6 named live. Now #4/#5/#6, and §6's Forks bullet — which still listed three — records the real inventory: six forks, five open. MAJOR (round 7) — the §1.7 note asserted the §7 S-I site "was found by the author's own class grep, not by any review naming it". False: round 6's watch-list names that exact line. The claim is retracted in place, and the note now states plainly that the author-side sweep never became class-driven in any round. MINORs: "no magnitude word is used here" was false of its own sentence; the §8 pointer omitted §8.6; "five rounds running" and "round 5's findings" were inherited prose that silently re-pointed to a different commit. Prior-art: skipped — review-absorption edit on existing research patches, no new capability
…asis says, repair the broken substitution Round 8 recorded the first CLEAN on W-5, the cross-population-share criterion that produced the MAJOR in rounds 4, 5, 6 and 7. The remaining findings are a different and narrower class. MAJOR — §8.4 stated its withdrawal broader than the §8.2 reading it cites as basis: §8.2 withdraws shares "of the current head", §8.4 restated that as "every share of it [the pre-S-G set] is withdrawn" — while four surfaces still publish exactly such shares (§8.2's class table, §8.6's four-denominator table and their two parent-side twins) and while holding #6 open, whose entire subject is which share to publish. Narrowed to match: the pre-S-G shares stay valid AS pre-S-G shares and are not withdrawn; what is withdrawn is any share of, or top lever for, the CURRENT set — which is what a subtraction map acts on. #6 is now named as the question of which denominator a pre-S-G share is quoted against. MAJOR — propagating the share-withdrawal into the §7 S-I surface broke the sentence: "the listing is nonetheless reduced to measured at 8.9k injected" left a dangling verb phrase, asserted 8.9k twice, and kept the direction word "reduced" that R4 forbids six lines into its own text. Rewritten to carry R4's own closing position: the budget premise is wrong by ~4x, and NO claim is made about truncation either way. MINORs: §1.7's marker inventory still read "R5 REVERSED" after this round renamed it SUSPENDED; the -20,782 B set cut was attributed entirely to the traps->digest swap, which accounts for -19,684 B (the rest is two other files in the same trim); §8.4 called §8.1's measured B/token aggregate an "identity" alongside §8.2's exact arithmetic one, upgrading a 2.37-3.32 empirical average to an exact relation in the round whose purpose was the opposite; "the file that dominates this measurement" is 9.8k against CLAUDE.md's 9.3k, so it is the largest single file, not a dominant one. Prior-art: skipped — review-absorption edit on existing research patches, no new capability
…o §1.7, drop the wrong-side share label Round 9 resolved five of round 8's six findings and returned one MAJOR of the same W-17 class at a site the previous commit did not reach, plus one MINOR in the wording it introduced. MAJOR — §8.4 narrowed its withdrawal to "any share of the CURRENT set", but the §1.7 T15 paragraph still carried the pre-round-8 breadth: "(§8.2 reading 1, whose share figures are withdrawn as pre-S-G)". The file therefore issued two incompatible instructions about the same table to the same consumer, and the §1.7 form also dropped the "of the current head" qualifier its cited basis carries (§8.2 reading 1). Restated to match that basis exactly: a pre-S-G measurement from which no share of the current head is derived. Enumerated every withdrawal statement across all three S-H patches (grep -n withdraw → 6 hits: addendum :69, :159, :290, :327; parent :262, :297); this was the sole over-broad survivor — :262 withdraws the injected-vs-source share, :297 withdraws the R5 reversal, both correctly scoped. MINOR — the exemption introduced last round read "remain valid as shares of that pre-S-G set". That is exact for §8.2, whose denominator IS the pre-S-G resident head (58.2k), but inverted for §8.6, where the pre-S-G block (26,700) is the NUMERATOR and the four denominators are seat totals — the relation the same paragraph states correctly two lines later. Replaced with a form true of both: computed from that pre-S-G measurement, each against the denominator its own table names. Both edits are subtractive/narrowing and introduce no figure, ratio or magnitude word — the strategy that first produced a CLEAN on W-5 at round 7. Prior-art: skipped — review-absorption edit on existing research patches, no new capability
This was referenced Aug 7, 2026
Merged
artyhoo
added a commit
that referenced
this pull request
Aug 7, 2026
…rable gate MET (#1267) Rev 7 left five sites asserting that S-L is unmerged, the loudest being `…-s-d-prime/kickoff.md:361` «The stage still cannot start, for one remaining reason: **S-L is not merged**». S-L merged as PR #1263 on 2026-08-07T12:50Z, so an executor reading the kickoff in full — which §0 requires — hits a stop-text that is now false. Dispatching against it would be `#dispatch-before-staging` in the other direction: the input on staging says «do not start». Retired at all five sites, three in the stage kickoff (`:1` header, `:310` §5 consequence line, `:361` dispatch status) and two in the umbrella (`:97` stage-table cell, `:362` S-L ordering paragraph). Scope: no deliverable, no permitted-file set, and no acceptance criterion changes. The rev-8 dispatch-status paragraph additionally carries forward the one binding thing S-L's §5 says about this stage — «A re-ranking is not a rescale — S-D′ must re-derive rather than multiply through» (`docs/meta-factory/research-patches/2026-08-07-s-l-recalculation.md:509-510`), plus the note that both hook injects are levers a `/context`-ordered list ranks at zero (`:499`) — because an executor that multiplies through a uniform factor preserves order by construction and would hide exactly the effect S-L found. Both gates verified mechanically, not from the umbrella prose (the recurrence that memory `verify-before-claim-family` records): S-E #1237 MERGED 09:39Z; meter present, scripts/measure-always-on.sh:10-11 S-H #1239 MERGED 00:06Z + #1249; P11 returned a real absence with a discriminating control, NOT INCONCLUSIVE (…-p3d-p11.md:431,440) S-L #1263 MERGED 12:50Z; §5 «Spec reach — what this does to the S-D′ ranking» (…-s-l-recalculation.md:493-510) Umbrella kickoff stays at 592 lines (600-line pre-commit gate): the two edits are in-place rewrites, no appends. Prior-art: skipped — doc-only correction to two dispatch-input kickoffs; retires a stale gate claim, adds no capability, no dependency, no code. Co-authored-by: Test <test@example.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Follow-up to the merged S-H stage PR #1239 (squash
cfe5ee34). After that merge the operatorpasted
/contextoutput, which is the measurement channel S-H'sDECISION-NEEDED #3asked for(Option A). This PR records it.
It lands as a new companion patch, not an in-place rewrite of the merged one — deliberately
(the merged patch is the record of what was knowable without the paste) and necessarily (the
parent is 496 lines against the repo's 600-line markdown gate). The addendum keeps
§8.xnumbering so the parent's cross-references resolve into it.
What the paste establishes:
wc -cbyte countand a harness token count: 77,156 B / 29,464 tok = 2.62 B/token. Every 4 B/t figure in both
S-H patches is low by ≈1.53×, and because the harness-remainder row is computed by difference
it is correspondingly high. Recorded as a correction owed, not made →
DECISION-NEEDED #4.Re-deriving it here would be the unmeasured re-derivation T-SH-A forbids.
category rows sum to 58.2k — an exact identity, not an estimate.
ToolSearchdeferral withholds58.1k, about the size of the whole resident head. It does not halve the budget (free space
moves −8.7% on this 1m window); it doubles the head, which every seat pays before its first
word.
UNMEASUREDrows close (5c MCP tool schemas 8.4k; 5e skills 8.9k + agents 1k).5d and row 9 stay
UNMEASURED — channel absent—/contextdoes not itemise MCP serverinstructions apart from tool schemas, and its «Custom agents» row counts registered agent types,
not the repo's
agents/directory. Reporting two of four where four was expected is T-SH-Aworking, not a shortfall.
channel defines «harness remainder» — the by-difference and
/contextchannels disagree by~30.8k on the same seat), feat(phase-7): L4 Validator + L5 Installer #6 (which of four defensible denominators the repo-owned share is
quoted against, which decides whether ADR-3's 29-39% band reads high, low or inside) are
surfaced with options and left to the operator per kickoff §3a. feat: audit fixes Phase 1+2+3 (R2/R11 drift, manifest SSOT, depcruise integration) #1 and feat: self-application + meta-factory foundation (Phase 0.5–3) #2 live in the sibling
patch.
Changes
docs/meta-factory/research-patches/2026-08-07-s-h-p14-context-addendum.md(new, 329 lines) —the
/contextmeasurement, its two readings, the closed and still-openUNMEASUREDrows, thethree new forks, and its own §1.7 backward sweep.
docs/meta-factory/research-patches/2026-08-07-s-h-harness-remainder-p14.md(edited) — theparent's sections whose evidentiary basis the paste moved now carry markers: §0 supersession,
§0a ANSWERED, §2 headline warning + per-row seat annotations, §4 R1 PERFORMED / R4 PARTLY
CHALLENGED / R5 SUSPENDED, §6 revised partition, §7 T-SH-A revision and all four backward-check
surfaces re-adjudicated.
Both files are inside the kickoff's §2 permitted set. Nothing else is touched:
git diff --name-only origin/staging...HEADreturns exactly these two paths.Prior-art consult
Prior-art:trailer — the escape hatch form. Not a capability commitunder the CLAUDE.md mechanical definition: no new dependency, no file under
packages/,no new
packages/core/<dir>/.Last reviewedtouch required.build-first-reuse-default.md §3toolingcaveat context7 targets library API docs, not this problem class.
Test plan
npx vitest run packages/core/principles/— 36 files / 348 passed, 1 skipped, at headcc080f31ba.npx vitest run packages/core/principles/10-research-patch-annotation.test.ts(kickoff §3contract line 4) — green; both patches carry their
<!-- scope:… -->first line.cc080f31ba(no--no-verify).scripts/measure-turn-attribution.sh) — N/A to this diff:the script merged with arch-v2: stage S-H — host-side measurements (P3d per-turn attribution SSOT + P11 Explore/Plan probe + P14 harness price list) #1239 and is untouched here.
Provenance
Kickoff
.claude/orchestrator-prompts/arch-v2-context-pipeline-s-h/kickoff.md(rev 5) · baseorigin/staging· substrate in-session (host CC) — structural per the kickoff header: the/contextslash command exists only in a live host session, and the aif container mountsclaude-authas a named volume rather than the host~/.claude· models: implementation Opus,cold fidelity audits Opus · fidelity Round 10.
The verbatim
/contextcapture this patch quotes is preserved outside the repo at~/.claude/projects/-Users-art-code-rules-as-tests-aif/context-capture-2026-08-07.md; it is thesingle source of every figure in §8.
Review findings
Ten cold fidelity rounds on this follow-up — nine REVISE, then GO. Every round was a fresh
seat that never saw the implementation session; rounds 2-10 were narrow-delta rounds handed the
incremental diff, the kickoff's scope sections and the Watch-list inlined, per
.claude/rules/cold-seat-economy.md§3 — never a resumed transcript. (Rounds 1-4 of the stagePR #1239 are a separate, closed series with its own watch-list; this counter starts at 1 for the
follow-up.)
The rounds are worth recording because they are one defect class, found nine times:
channel partition (W-2), and a false population-mismatch escape used to avoid a correction (W-9).
round's own replacement figure. The invariant: a quantity derived across mismatched
populations, denominators or conversion constants — or a word substituted for a number that
states a different claim than the number did. Round 6's replacement inverted a 56-87% share into
«a minority». The two concrete traps, now binding on anyone editing these patches: the harness's
total for the 74 listed skill entries is not divisible by a byte sum over 129
SKILL.mdfiles (
datavizandclaude-api, the two largest listed entries, have noSKILL.mdat all),and 26,700 is the pre-S-G resident set, so no ranking of the current set may be built on it.
arithmetic was the failure mode.
were a narrower class: a withdrawal stated broader than the section it cited as basis (W-17), and
a sentence broken by the substitution that left the withdrawn direction word standing (W-16).
re-derived by the auditor. One MAJOR remained, and notably not in the new wording: a stale
§1.7 sentence still carrying the pre-round-8 breadth, which the fix had not reached. Plus one
MINOR in the new exemption label, which put the pre-S-G set on the denominator side where §8.6
has it as the numerator → W-18.
the full population rather than sampling:
grep -n "withdraw"→ 6 hits across both patches, allcorrectly scoped;
grep -n "pre-S-G"→ 10 hits, every one read. No surface in either patch stillstates the withdrawal at the pre-round-8 breadth. One non-blocking MINOR opened as W-19 and
carried rather than fixed, so it reaches S-D′ as a recorded item with its reintroduction tell.
Watch-list
/contextrow the capture lacksp14.md:264-270,:462-464; addendum:112-118,:146-149;p14.md:298-309(r7)git diff --name-onlyrow outsidescripts/measure-turn-attribution.sh+research-patches/*:266(r8)findas a correction repeats the error in the other direction:282-288(r7):150-151;p14.md:362-364,:486-489(r7):259,:265-266(r7)p14.md:266(r7);p14.md:462-463(r8):154-157(r8); addendum:327(r9):157-158(r9):157-159(live r10)Round 8: W-1 CLEAN · W-2 CLEAN · W-3 CLEAN · W-4 CLEAN · W-5 CLEAN · W-6 CLEAN · W-7 REINTRODUCED (addendum:266) · W-8 CLEAN · W-9 CLEAN · W-10 CLEAN · W-11 CLEAN · W-12 CLEAN · W-13 CLEAN · W-14 CLEAN · W-15 CLEAN · W-16 REINTRODUCED (p14.md:462-463) · W-17 REINTRODUCED (addendum:154-157)
Round 9: W-1 CLEAN · W-2 CLEAN · W-3 CLEAN · W-4 CLEAN · W-5 CLEAN · W-6 CLEAN · W-7 CLEAN · W-8 CLEAN · W-9 CLEAN · W-10 CLEAN · W-11 CLEAN · W-12 CLEAN · W-13 CLEAN · W-14 CLEAN · W-15 CLEAN · W-16 CLEAN · W-17 REINTRODUCED (addendum:327) · W-18 NEW (addendum:157-158)
Round 10: W-1 CLEAN · W-2 CLEAN · W-3 CLEAN · W-4 CLEAN · W-5 CLEAN · W-6 CLEAN · W-7 CLEAN · W-8 CLEAN · W-9 CLEAN · W-10 CLEAN · W-11 CLEAN · W-12 CLEAN · W-13 CLEAN · W-14 CLEAN · W-15 CLEAN · W-16 CLEAN · W-17 CLEAN · W-18 CLEAN · W-19 NEW (addendum:157-159)
Fidelity verdict
FIDELITY: GO
Basis: .claude/orchestrator-prompts/arch-v2-context-pipeline-s-h/kickoff.md#§3 + §3a + §4
Round: 10
Audited-SHA: cc080f3
Evidence: docs/meta-factory/research-patches/2026-08-07-s-h-p14-context-addendum.md:326-327 now states exactly what its cited basis at :119-120 states, and :158-159 places the pre-S-G set on the numerator side, matching §8.6's table at :219-224 and the parent-side twin at docs/meta-factory/research-patches/2026-08-07-s-h-harness-remainder-p14.md:429-434;
grep -n "withdraw"over both patches returns six hits, all correctly scoped;grep -n "pre-S-G"returns ten, every one read — no surface still states the withdrawal at the pre-round-8 breadth.Findings:
wc -c× 4 B/t-derived, so a reader binding it into that set would read the wrong channel — the T-SH-A dimension. Carried as W-19 rather than fixed: the natural referent is exact, no figure, ratio, direction word or misplaced numerator is introduced, and a further edit would re-open a ten-round loop for a referent the next reader of §8.4 resolves correctly.§1.7 Forward-check applied
.claude/rules/attention-is-not-a-mechanism.md:15(§1) — the acceptance layer here is branch(b), a NAMED cold-agent protocol with structured output (
agents/fidelity-auditor.md:74, theverdict rule), transported by the fail-closed
pr-body-fidelityrequired check(
packages/core/hooks/checks/pr-body-fidelity.ts:165enforces theAudited-SHA-prefixes-headmatch). Not «a reviewer will read the diff». Nine of ten rounds returned REVISE, which is the
evidence that the detection layer does work rather than ratify.
.claude/rules/ai-laziness-traps.md:56(T3),:122(T14),:159(T20) + the kickoff'sT-SH-A — every figure in the addendum carries its channel (
/contextrow, orwc -c); the twoblocks with no channel keep
UNMEASURED — channel absentrather than a plausible neighbour'snumber; the conversion correction the paste makes necessary is recorded as owed, not made,
because making it would be an unmeasured re-derivation. Coverage is stated as predicates (two of
four rows closed), never as «high confidence» or as «clean».
.claude/rules/cold-seat-economy.md:56(§3) — rounds 2-10 were fresh narrow seats with thewatch-list inlined, never resumed transcripts (
:109#continuity-by-replay), and no round'sverdict was self-issued by the editing session (
:102#self-issued-verdict)..claude/rules/no-paid-llm-in-ci.md:20(§1) — every audit is a session-read agent; nothingadded to CI.
.claude/rules/language-discipline.md:20(§1) — machinery and artefacts in English..claude/rules/build-first-reuse-default.md:66— REUSE: no new artefact beyond the researchpatch itself, and context7 is correctly not consulted (its tooling caveat scopes it to library
API docs, not this problem class).
§1.7 Backward-check applied
Class of this change = a superseding measurement that moves the evidentiary basis of already
published figures, and the withdrawal/exemption statements that follow from it. Two enumerations,
both mechanical:
docs/meta-factory/research-patches/2026-08-07-s-h-p14-context-addendum.md:294-319, verdictedper surface: the parent patch SWEPT; the sibling
2026-08-07-s-h-turn-attribution-p3d-p11.mdGAP-FOUND, not edited (it carries the same 4 B/tconstant at §5/§7/§8, so §8.1's falsification applies identically — but re-deriving it is the
re-measurement Phase 4: Stack Detector v1 (read+write AIF bridge, CI gate, reviewer cycle complete) #4 must settle first, and it is named in Phase 4: Stack Detector v1 (read+write AIF bridge, CI gate, reviewer cycle complete) #4's Option A as required scope);
docs/superpowers/specs/2026-08-06-pipeline-token-economy-design.mdand…/2026-07-31-arch-v2-context-pipeline-design.mdADR-3 GAP-FOUND, out of permitted set(spec-level, round-capped, operator-owned — surfaced via Phase 4: Stack Detector v1 (read+write AIF bridge, CI gate, reviewer cycle complete) #4, not edited);
2026-08-01-token-economy-s-a-profile.mdNOT SWEPT by ownership;.claude/settings.jsonNOT SWEPT, deliberately (operator-only, agent-uncommittable).
grep -n "withdraw"over bothpatches returns exactly six sites (addendum
:69,:159,:290,:327; parent:262,:297) andgrep -n "pre-S-G"ten; all sixteen were read, not sampled. Five of the sixwithdrawals were already scoped to a specific claim or to the current set; the sixth
(addendum
:327) was the round-9 MAJOR and is fixed atcc080f31ba. The parent's two arecorrect as they stand:
:262withdraws the injected-vs-source skills share,:297withdrawsthe R5 reversal while explicitly limiting the pre-S-G snapshot to «cannot rank the current
set».
No surface is left inconsistent, and no surface outside the permitted set was edited — the three
GAP-FOUND ones are routed with the fork that must settle them first.
Parked questions
Three
DECISION-NEEDEDforks are surfaced per kickoff §3a and left unresolved — they are theoperator's. All three live in the patch itself, so they survive independently of this PR body.
files that carry both counts. Option A — re-convert every
wc -c-derived figure in both S-Hpatches (named scope includes the sibling patch and the spec's
[W]/[H]rows). Option B —leave the figures and carry the ≈1.53× correction as a standing caveat. Doing nothing leaves the
by-difference harness remainder biased high by a known factor.
/contextdisagree by ~30.8k on the same seat. Until one is named operative, «68.4% of a subagent seat»
and the
/contextsplit are two different quantities wearing one name.(89,019 / 100,529 / 58,200 / 62,340), disagreeing in direction against ADR-3's 29-39% band:
inside, below, above, above. feat(phase-7): L4 Validator + L5 Installer #6 cannot be settled independently of feat(meta-factory): Phase 5+6 — L2 Research Agent + L3 Synthesizer Path A (deterministic v1) #5.
Observation (no PR spawned, per CLAUDE.md
PR strategy)Ten rounds went to one discipline that is general, mechanically checkable at review time, and not
currently codified anywhere: a share's numerator must be provably a subset of its denominator,
and a word substituted for a withdrawn figure must keep that figure's direction. W-5, W-13,
W-16, W-17, W-18 and W-19 are all specialisations of it. It looks like a
.claude/rules/candidate. Surfaced as an observation, not acted on — the umbrella scope is S-H.