Skip to content

docs(s-h): P14 /context addendum — the operator paste that answers DECISION-NEEDED #3 - #1249

Merged
artyhoo merged 13 commits into
stagingfrom
worktree-s-h-p14
Aug 7, 2026
Merged

docs(s-h): P14 /context addendum — the operator paste that answers DECISION-NEEDED #3#1249
artyhoo merged 13 commits into
stagingfrom
worktree-s-h-p14

Conversation

@artyhoo

@artyhoo artyhoo commented Aug 7, 2026

Copy link
Copy Markdown
Owner

Summary

Follow-up to the merged S-H stage PR #1239 (squash cfe5ee34). After that merge the operator
pasted /context output, which is the measurement channel S-H's DECISION-NEEDED #3 asked for
(Option A). This PR records it.

It lands as a new companion patch, not an in-place rewrite of the merged one — deliberately
(the merged patch is the record of what was knowable without the paste) and necessarily (the
parent is 496 lines against the repo's 600-line markdown gate). The addendum keeps §8.x
numbering so the parent's cross-references resolve into it.

What the paste establishes:

Changes

  • docs/meta-factory/research-patches/2026-08-07-s-h-p14-context-addendum.md (new, 329 lines) —
    the /context measurement, its two readings, the closed and still-open UNMEASURED rows, the
    three new forks, and its own §1.7 backward sweep.
  • docs/meta-factory/research-patches/2026-08-07-s-h-harness-remainder-p14.md (edited) — the
    parent's sections whose evidentiary basis the paste moved now carry markers: §0 supersession,
    §0a ANSWERED, §2 headline warning + per-row seat annotations, §4 R1 PERFORMED / R4 PARTLY
    CHALLENGED / R5 SUSPENDED, §6 revised partition, §7 T-SH-A revision and all four backward-check
    surfaces re-adjudicated.

Both files are inside the kickoff's §2 permitted set. Nothing else is touched:
git diff --name-only origin/staging...HEAD returns exactly these two paths.

Prior-art consult

  • Every commit carries a Prior-art: trailer — the escape hatch form. Not a capability commit
    under the CLAUDE.md mechanical definition: no new dependency, no file under packages/,
    no new packages/core/<dir>/.
  • No new capability area surfaced → no new SSOT entry required.
  • No existing SSOT entry matched → no Last reviewed touch required.
  • context7 N/A — no new capability area, and per build-first-reuse-default.md §3 tooling
    caveat context7 targets library API docs, not this problem class.

Test plan

  • npx vitest run packages/core/principles/36 files / 348 passed, 1 skipped, at head
    cc080f31ba.
  • npx vitest run packages/core/principles/10-research-patch-annotation.test.ts (kickoff §3
    contract line 4) — green; both patches carry their <!-- scope:… --> first line.
  • markdownlint — 0 errors on the edited file.
  • 600-line markdown gate — 329 / 496 lines, both under.
  • Pre-push hook ran clean on the push of cc080f31ba (no --no-verify).
  • Kickoff §3 contract lines 1-3 (scripts/measure-turn-attribution.sh) — N/A to this diff:
    the script merged with arch-v2: stage S-H — host-side measurements (P3d per-turn attribution SSOT + P11 Explore/Plan probe + P14 harness price list) #1239 and is untouched here.

Provenance

Kickoff .claude/orchestrator-prompts/arch-v2-context-pipeline-s-h/kickoff.md (rev 5) · base
origin/staging · substrate in-session (host CC) — structural per the kickoff header: the
/context slash command exists only in a live host session, and the aif container mounts
claude-auth as a named volume rather than the host ~/.claude · models: implementation Opus,
cold fidelity audits Opus · fidelity Round 10.

The verbatim /context capture this patch quotes is preserved outside the repo at
~/.claude/projects/-Users-art-code-rules-as-tests-aif/context-capture-2026-08-07.md; it is the
single source of every figure in §8.

Review findings

Ten cold fidelity rounds on this follow-up — nine REVISE, then GO. Every round was a fresh
seat that never saw the implementation session; rounds 2-10 were narrow-delta rounds handed the
incremental diff, the kickoff's scope sections and the Watch-list inlined, per
.claude/rules/cold-seat-economy.md §3 — never a resumed transcript. (Rounds 1-4 of the stage
PR #1239 are a separate, closed series with its own watch-list; this counter starts at 1 for the
follow-up.)

The rounds are worth recording because they are one defect class, found nine times:

  • Rounds 1-3 — the paste's figures stated wider than the snapshot supports (W-1), a stale
    channel partition (W-2), and a false population-mismatch escape used to avoid a correction (W-9).
  • Rounds 4-7 — the load-bearing stretch. Each round found its MAJOR inside the previous
    round's own replacement figure
    .
    The invariant: a quantity derived across mismatched
    populations, denominators or conversion constants — or a word substituted for a number that
    states a different claim than the number did. Round 6's replacement inverted a 56-87% share into
    «a minority». The two concrete traps, now binding on anyone editing these patches: the harness's
    total for the 74 listed skill entries is not divisible by a byte sum over 129 SKILL.md
    files
    (dataviz and claude-api, the two largest listed entries, have no SKILL.md at all),
    and 26,700 is the pre-S-G resident set, so no ranking of the current set may be built on it.
  • Round 7 changed strategy: stop deriving, start withdrawing. Removal needs no arithmetic, and
    arithmetic was the failure mode.
  • Round 8 — first CLEAN on W-5, the cross-population criterion. Its two remaining MAJORs
    were a narrower class: a withdrawal stated broader than the section it cited as basis (W-17), and
    a sentence broken by the substitution that left the withdrawn direction word standing (W-16).
  • Round 9 — five of round 8's six findings resolved with the arithmetic independently
    re-derived by the auditor. One MAJOR remained, and notably not in the new wording: a stale
    §1.7 sentence still carrying the pre-round-8 breadth, which the fix had not reached. Plus one
    MINOR in the new exemption label, which put the pre-S-G set on the denominator side where §8.6
    has it as the numerator → W-18.
  • Round 10 — GO. Both round-9 findings resolved; all 18 rows CLEAN. The auditor enumerated
    the full population rather than sampling: grep -n "withdraw" → 6 hits across both patches, all
    correctly scoped; grep -n "pre-S-G" → 10 hits, every one read. No surface in either patch still
    states the withdrawal at the pre-round-8 breadth. One non-blocking MINOR opened as W-19 and
    carried rather than fixed, so it reaches S-D′ as a recorded item with its reintroduction tell.

Watch-list

id criterion why defect site reintroduction tell
W-1 a quantitative claim states only what the quoted snapshot measures the snapshot is n=1 on one seat; anything wider is an estimate T-SH-A bans fixed r2 a figure quoted without its seat/channel qualifier
W-2 price-row channel counts match the table the partition is the T-SH-A compliance evidence; a stale count fakes it fixed r2-3 a partition string that does not survive a hand recount of the two §2 tables
W-3 T-SH-A restraint rows stay unfilled, basis matches the capture filling 5d/9 from a neighbour is the estimate-as-measurement failure fixed r2 5d or 9 gaining a number, or a basis cell citing a /context row the capture lacks
W-4 the conversion correction is recorded as owed, not made silently reconverting §2 would republish an unmeasured re-derivation none — preventive any §2 figure changing without a stated re-measurement
W-5 figures from different seats/denominators/populations are never divided by or summed against each other; a share's numerator must be provably a subset of its denominator the recurring MAJOR class across rounds 4-7 p14.md:264-270, :462-464; addendum :112-118, :146-149; p14.md:298-309 (r7) any new %, ratio or "×" whose two sides are not shown to share a population
W-6 scope — permitted set only kickoff §2 lists the two research patches + the script; anything else is scope creep none — preventive a git diff --name-only row outside scripts/measure-turn-attribution.sh + research-patches/*
W-7 forks surfaced as DECISION-NEEDED; §1.7 verdicts and marker names re-adjudicated when the same commit restates them a §1.7 record that misnames a verdict the commit changed misleads the next reader about what was swept addendum :266 (r8) a §1.7 bullet naming a marker string absent from the parent
W-8 the §4 sweep reaches EVERY recommendation whose basis the paste moved R1/R2/R4/R5 all rest on the paste; a missed one ships a superseded action fixed r4 an R-block with no post-paste marker while its evidence moved
W-9 a caveat justifying NOT correcting a figure is itself verified before being called binding r3 caught a false population-mismatch escape used to avoid a correction fixed r7 a "not corrected because X" whose X is asserted, not shown
W-10 before «X is essentially Y», both sides share a population AND a constant, stated the equivalence hides the same division defect W-5 catches fixed r6 "essentially", "almost exactly", "roughly the same as" between two channels
W-11 a corrected population must be the command's output AFTER filtering to what the claim is about publishing a raw find as a correction repeats the error in the other direction fixed r7 a population offered as a fix whose command includes node_modules/worktrees/cache
W-12 a §1.7 note naming a recurring method failure either applies the counter it cites or states that the current round still exhibits it naming T21 while committing it is the trap itself addendum :282-288 (r7) a §1.7 claiming the sweep became class-driven, or omitting the current round from the failure
W-13 a figure withdrawn to a qualitative word must keep the side of the numbers it replaces r6 replaced a majority with "a minority" fixed r7 a qualitative word whose direction contradicts the withdrawn figure
W-14 a newly raised DECISION-NEEDED propagates to every surface that enumerates the fork set a stale fork count reads as a settled question addendum :150-151; p14.md:362-364, :486-489 (r7) a fork enumeration stopping before the highest-numbered open fork
W-15 a round's self-references («this round», «N rounds running») are re-anchored when a later round edits the same bullet a frozen round count dates the note and misstates the record addendum :259, :265-266 (r7) a hard round count or "round N's findings" in a bullet later rounds touch
W-16 when a justification is removed, the figures and the direction words it supported are re-derived or narrowed — not left standing the removal is the whole point; a surviving claim re-publishes what was withdrawn p14.md:266 (r7); p14.md:462-463 (r8) a clause reading "…is nonetheless <direction word>" after its supporting figure was withdrawn, or a verb phrase left dangling by the substitution
W-17 a withdrawal states no more than the section it cites as its basis, and every surface still publishing what it withdraws is either updated or explicitly exempted an over-broad withdrawal silently voids live tables and open forks elsewhere in the same document addendum :154-157 (r8); addendum :327 (r9) "every share/ranking/figure … is withdrawn" without the qualifier its cited basis carries, while a share table or a DECISION-NEEDED over that same quantity stays published
W-18 a prose label naming what a published share is a share of must place the named set on the side (numerator or denominator) it actually occupies in the cited table r9 introduced the exemption mechanism ("these shares stay valid as shares of X"), which future rounds will restate; naming the wrong side re-describes a live table as measuring something it does not addendum :157-158 (r9) an exemption or withdrawal phrased "shares of <set>" where the cited table's denominator is not that set
W-19 a provenance verb ("computed from", "measured by", "sourced from") applies only to surfaces the sentence enumerates; a collective noun standing in for unnamed sibling tables must be resolvable to tables that actually share that channel r10 replaced a side-label with a channel claim; the channel is the T-SH-A compliance dimension, and a collective referent can silently attach the wrong channel to a row measured another way addendum :157-159 (live r10) "and their parent-side twins / the corresponding rows / the same figures elsewhere" governed by a single channel or provenance verb, where at least one member of the implied set is measured on a different channel

Round 8: W-1 CLEAN · W-2 CLEAN · W-3 CLEAN · W-4 CLEAN · W-5 CLEAN · W-6 CLEAN · W-7 REINTRODUCED (addendum:266) · W-8 CLEAN · W-9 CLEAN · W-10 CLEAN · W-11 CLEAN · W-12 CLEAN · W-13 CLEAN · W-14 CLEAN · W-15 CLEAN · W-16 REINTRODUCED (p14.md:462-463) · W-17 REINTRODUCED (addendum:154-157)
Round 9: W-1 CLEAN · W-2 CLEAN · W-3 CLEAN · W-4 CLEAN · W-5 CLEAN · W-6 CLEAN · W-7 CLEAN · W-8 CLEAN · W-9 CLEAN · W-10 CLEAN · W-11 CLEAN · W-12 CLEAN · W-13 CLEAN · W-14 CLEAN · W-15 CLEAN · W-16 CLEAN · W-17 REINTRODUCED (addendum:327) · W-18 NEW (addendum:157-158)
Round 10: W-1 CLEAN · W-2 CLEAN · W-3 CLEAN · W-4 CLEAN · W-5 CLEAN · W-6 CLEAN · W-7 CLEAN · W-8 CLEAN · W-9 CLEAN · W-10 CLEAN · W-11 CLEAN · W-12 CLEAN · W-13 CLEAN · W-14 CLEAN · W-15 CLEAN · W-16 CLEAN · W-17 CLEAN · W-18 CLEAN · W-19 NEW (addendum:157-159)

Fidelity verdict

FIDELITY: GO
Basis: .claude/orchestrator-prompts/arch-v2-context-pipeline-s-h/kickoff.md#§3 + §3a + §4
Round: 10
Audited-SHA: cc080f3
Evidence: docs/meta-factory/research-patches/2026-08-07-s-h-p14-context-addendum.md:326-327 now states exactly what its cited basis at :119-120 states, and :158-159 places the pre-S-G set on the numerator side, matching §8.6's table at :219-224 and the parent-side twin at docs/meta-factory/research-patches/2026-08-07-s-h-harness-remainder-p14.md:429-434; grep -n "withdraw" over both patches returns six hits, all correctly scoped; grep -n "pre-S-G" returns ten, every one read — no surface still states the withdrawal at the pre-round-8 breadth.
Findings:

  • MINOR docs/meta-factory/research-patches/2026-08-07-s-h-p14-context-addendum.md:157-159 — the provenance verb "computed from that pre-S-G measurement" governs an unenumerated "their parent-side twins". The twin that publishes shares (2026-08-07-s-h-harness-remainder-p14.md:429-434) matches exactly, since its numerator is the same 26,700 sourced to addendum §8.2; but p14.md:141's 27.8% is wc -c × 4 B/t-derived, so a reader binding it into that set would read the wrong channel — the T-SH-A dimension. Carried as W-19 rather than fixed: the natural referent is exact, no figure, ratio, direction word or misplaced numerator is introduced, and a further edit would re-open a ten-round loop for a referent the next reader of §8.4 resolves correctly.

§1.7 Forward-check applied

  • .claude/rules/attention-is-not-a-mechanism.md:15 (§1) — the acceptance layer here is branch
    (b), a NAMED cold-agent protocol with structured output (agents/fidelity-auditor.md:74, the
    verdict rule), transported by the fail-closed pr-body-fidelity required check
    (packages/core/hooks/checks/pr-body-fidelity.ts:165 enforces the Audited-SHA-prefixes-head
    match). Not «a reviewer will read the diff». Nine of ten rounds returned REVISE, which is the
    evidence that the detection layer does work rather than ratify.
  • .claude/rules/ai-laziness-traps.md:56 (T3), :122 (T14), :159 (T20) + the kickoff's
    T-SH-A
    — every figure in the addendum carries its channel (/context row, or wc -c); the two
    blocks with no channel keep UNMEASURED — channel absent rather than a plausible neighbour's
    number; the conversion correction the paste makes necessary is recorded as owed, not made,
    because making it would be an unmeasured re-derivation. Coverage is stated as predicates (two of
    four rows closed), never as «high confidence» or as «clean».
  • .claude/rules/cold-seat-economy.md:56 (§3) — rounds 2-10 were fresh narrow seats with the
    watch-list inlined, never resumed transcripts (:109 #continuity-by-replay), and no round's
    verdict was self-issued by the editing session (:102 #self-issued-verdict).
  • .claude/rules/no-paid-llm-in-ci.md:20 (§1) — every audit is a session-read agent; nothing
    added to CI.
  • .claude/rules/language-discipline.md:20 (§1) — machinery and artefacts in English.
  • .claude/rules/build-first-reuse-default.md:66 — REUSE: no new artefact beyond the research
    patch itself, and context7 is correctly not consulted (its tooling caveat scopes it to library
    API docs, not this problem class).

§1.7 Backward-check applied

Class of this change = a superseding measurement that moves the evidentiary basis of already
published figures, and the withdrawal/exemption statements that follow from it.
Two enumerations,
both mechanical:

  1. Documents whose basis this paste moves — enumerated in the addendum's own §1.7 at
    docs/meta-factory/research-patches/2026-08-07-s-h-p14-context-addendum.md:294-319, verdicted
    per surface: the parent patch SWEPT; the sibling
    2026-08-07-s-h-turn-attribution-p3d-p11.md GAP-FOUND, not edited (it carries the same 4 B/t
    constant at §5/§7/§8, so §8.1's falsification applies identically — but re-deriving it is the
    re-measurement Phase 4: Stack Detector v1 (read+write AIF bridge, CI gate, reviewer cycle complete) #4 must settle first, and it is named in Phase 4: Stack Detector v1 (read+write AIF bridge, CI gate, reviewer cycle complete) #4's Option A as required scope);
    docs/superpowers/specs/2026-08-06-pipeline-token-economy-design.md and
    …/2026-07-31-arch-v2-context-pipeline-design.md ADR-3 GAP-FOUND, out of permitted set
    (spec-level, round-capped, operator-owned — surfaced via Phase 4: Stack Detector v1 (read+write AIF bridge, CI gate, reviewer cycle complete) #4, not edited);
    2026-08-01-token-economy-s-a-profile.md NOT SWEPT by ownership; .claude/settings.json
    NOT SWEPT, deliberately (operator-only, agent-uncommittable).
  2. Every withdrawal/exemption statement in the permitted setgrep -n "withdraw" over both
    patches returns exactly six sites (addendum :69, :159, :290, :327; parent :262,
    :297) and grep -n "pre-S-G" ten; all sixteen were read, not sampled. Five of the six
    withdrawals were already scoped to a specific claim or to the current set; the sixth
    (addendum :327) was the round-9 MAJOR and is fixed at cc080f31ba. The parent's two are
    correct as they stand: :262 withdraws the injected-vs-source skills share, :297 withdraws
    the R5 reversal while explicitly limiting the pre-S-G snapshot to «cannot rank the current
    set».

No surface is left inconsistent, and no surface outside the permitted set was edited — the three
GAP-FOUND ones are routed with the fork that must settle them first.

Parked questions

Three DECISION-NEEDED forks are surfaced per kickoff §3a and left unresolved — they are the
operator's. All three live in the patch itself, so they survive independently of this PR body.

  1. Phase 4: Stack Detector v1 (read+write AIF bridge, CI gate, reviewer cycle complete) #4 — the conversion constant. The seed's 4 B/token is falsified at 2.62 B/token on seven
    files that carry both counts. Option A — re-convert every wc -c-derived figure in both S-H
    patches (named scope includes the sibling patch and the spec's [W]/[H] rows). Option B
    leave the figures and carry the ≈1.53× correction as a standing caveat. Doing nothing leaves the
    by-difference harness remainder biased high by a known factor.
  2. feat(meta-factory): Phase 5+6 — L2 Research Agent + L3 Synthesizer Path A (deterministic v1) #5 — which channel defines «harness remainder». The by-difference channel and /context
    disagree by ~30.8k on the same seat. Until one is named operative, «68.4% of a subagent seat»
    and the /context split are two different quantities wearing one name.
  3. feat(phase-7): L4 Validator + L5 Installer #6 — which denominator the repo-owned share is quoted against. Four defensible ones
    (89,019 / 100,529 / 58,200 / 62,340), disagreeing in direction against ADR-3's 29-39% band:
    inside, below, above, above. feat(phase-7): L4 Validator + L5 Installer #6 cannot be settled independently of feat(meta-factory): Phase 5+6 — L2 Research Agent + L3 Synthesizer Path A (deterministic v1) #5.

Observation (no PR spawned, per CLAUDE.md PR strategy)

Ten rounds went to one discipline that is general, mechanically checkable at review time, and not
currently codified anywhere: a share's numerator must be provably a subset of its denominator,
and a word substituted for a withdrawn figure must keep that figure's direction.
W-5, W-13,
W-16, W-17, W-18 and W-19 are all specialisations of it. It looks like a .claude/rules/
candidate. Surfaced as an observation, not acted on — the umbrella scope is S-H.

Test added 13 commits August 7, 2026 03:22
…EDED #3, and falsifies the 4 B/token convention

The operator ran `/context` post-merge and supplied the output, taking Option A of
DECISION-NEEDED #3 (§0a). Recorded as a new §8 addendum rather than an in-place rewrite: the
measurement history must read "unknown at stage close -> known 2026-08-07", not as though the
split had been available all along. §0a is kept verbatim, annotated ANSWERED.

Two findings, in order of consequence:

§8.1 — the seed's binding 4 B ~ 1 token conversion is FALSIFIED. Seven files carry both a
`wc -c` byte count and a harness-reported token count; aggregate 77,156 B / 29,464 tok =
2.62 B/token. Every 4 B/t figure in this patch and its sibling is low by ~1.53x, and because
row 5 (harness remainder) is computed BY DIFFERENCE, the remainder is correspondingly HIGH — a
first-order restatement puts it near 52%, not 68.4%. Figures are left as published and the
correction is recorded as owed, not made: the row-1 file set is the pre-S-G resident set while
the ratio was measured on the current one, so they are not the same population, and re-deriving
§2 on a new constant is a re-measurement beyond an addendum. Raised as DECISION-NEEDED #4.

§8.2 — the reported percentages sum to 105.6% because the two `(deferred)` rows are counted but
NOT resident. The identity confirms it exactly: 334.6k - 276.4k = 58.2k resident, and the
non-deferred rows sum to 58.2k. Half the resident head is memory files (29.4k of 58.2k), of
which two documents carry a third of everything (repo CLAUDE.md 9.3k + ai-laziness-traps.md
9.8k). ToolSearch deferral withholds 58.1k — almost exactly what the entire resident head
costs, which is the number §3's "preserve what already works" lacked.

§8.3 — closes TWO of the four `UNMEASURED — channel absent` rows, not four: 5c (MCP tool
schemas, 8.4k) and 5e (skills 8.9k + custom-agent listing 1k). 5d stays open (`/context` does
not itemise server instructions apart from tool schemas) and row 9 stays open (a different
population: "Custom agents" counts registered agent types, not the repo's agents/ directory).
Neither was filled from the nearest plausible neighbour — that is T-SH-A working, not a
shortfall. Revised partition 14 / 11 / 2 / 1, counted from the table.

All count-claims re-swept by class after the edit rather than site-by-site (the W-9 lesson from
the round-3 fidelity audit): table recount gives 14 rows and exactly 2 carrying the literal
marker; the stage-close claims of "four" are retained as historical and each carries its
revision inline.

Coverage: n=1, an orchestrator seat in a worktree with five rule files injected; a fresh
main-checkout or subagent seat has a different resident set. All figures are the harness's own
estimates at its own rounding; no tokenizer was run.

Prior-art: skipped — post-merge measurement addendum to an existing research patch, no new capability
MAJOR — §8.2 claimed `ToolSearch` deferral "roughly doubles the usable budget". The snapshot
cannot support that: window 1m, free space 665.4k, so making the 58.1k deferred schemas resident
moves free space to ~607.3k (-8.7%). What doubles is the resident HEAD (58.2k -> 116.3k).
Restated to the measure the snapshot actually bounds; the supported neighbouring claims (58.1k
is about the size of the whole head; still the most expensive available regression) are kept.

MINORs, all from the same cold seat:
- §0a heading was present-tense "five blocks stay unpriced", false after the update -> marked
  "(as at stage close) … stayed", with the current count (three: 5d, 9, row 8's injected form)
  stated in the ANSWERED block and again in §7.
- "five rule files" contradicted the patch's own table -> four, with the four named and the
  other three memory files identified.
- DECISION-NEEDED #4 Option A pointed at the sibling's "§5/§9"; the sibling has no §9 (it runs
  §0-§8) -> corrected to its actual 4 B/t sites, §5, §7 and §8.
- The 5d basis asserted server instructions "sit inside the system-prompt region"; the capture
  establishes only that /context does not itemise them apart from tool schemas -> the locational
  claim is dropped, since asserting a region is the estimate T-SH-A forbids.
- rows 1-4 restatement read 30,163, which reproduces from neither derivation route ->
  19,719 × (4/2.6187) = 30,120, remainder 62,340 - 30,120 = 32,220, share 51.7%.

All count-claims re-swept by class after the edit: table holds 14 rows with exactly 2 carrying
the literal UNMEASURED marker; every surviving "five" is either historical-and-marked or refers
to the item-4 probe's five files, a different subject.

Prior-art: skipped — review-absorption edit on an existing research patch, no new capability
…check, surface the channel disagreement

Round 2 confirmed all six round-1 findings closed and re-derived every §8 figure independently,
then found two MAJORs the addendum had not noticed about its own effect on the rest of the file.

MAJOR 1 — the §1.7 backward-check asserted SWEPT-CLEAN using figures this same commit restates.
Both verdicts re-adjudicated in place rather than left standing:
- ADR-3: the "inside ADR-3's stated band" clause was wrong when written — the band is 29-39%
  and both measurements (27.8% / ~21%) fall BELOW it; under §8.1's conversion the same share
  moves to ~47%, outside on the high side. Now GAP-FOUND, direction unresolved pending #4.
- the spec's P14 row: "the row's arithmetic holds" is true only under the 4 B/t constant it was
  computed with, since §8.1 restates the same seat at 51.7%. Now HOLDS-CONDITIONALLY on #4B.

MAJOR 2 — one seat, two irreconcilable harness figures, previously unflagged. By difference the
main seat's remainder is 69,300 of 89,019; /context's categories matching row 5's own definition
sum to 28.8k for that SAME session, and neither 28.8k nor 86.9k (adding deferred schemas back)
reaches 69,300. The totals disagree the same way: 58.2k resident vs 89,019 first-turn billed,
gap ~30.8k. New §8.5 states the disagreement, offers the dispatch-prompt hypothesis explicitly
as unmeasured (§0 defines the channel as "resident head PLUS its dispatch prompt", and rows 1-4
never subtract it; this session opened with /orchestrator, which injects a whole SKILL.md body),
and draws the consequence that matters: by-difference systematically OVERSTATES the remainder,
because anything it cannot attribute to rows 1-4 lands in row 5 by construction. Raised as
DECISION-NEEDED #5 with three options including "measure the gap directly". Not resolved here.

MINORs:
- §4 was the only section the revision sweep had skipped. R1 now carries a PERFORMED block (the
  paste happened; two of four rows closed, not four; S-D′ no longer has to park). R5's
  conclusion is REVERSED with its reasoning shown — its "next lever is harness-side" is
  contradicted by memory files being 50.5% of the resident head and repo-owned.
- rows 5c/5e now carry the seat annotation: orchestrator MAIN seat, n=1, not the 62,340-tok
  subagent seat the table is sized against, with an explicit do-not-sum-against-row-5.
- The headline now warns that both its percentages are contested, naming #4 and #5.

Count-claims re-swept: 14 table rows, exactly 2 carrying the literal UNMEASURED marker.

Prior-art: skipped — review-absorption edit on an existing research patch, no new capability
… the 600-line gate

Round 3 confirmed round-2's MAJOR #2 (channel disagreement) and MINOR #4 (seat annotations)
fully discharged, and re-derived every §8 figure independently. It then caught the replacement
figures themselves.

MAJOR — the ADR-3 re-verdict swapped one unsupported number for another: `29,464 / 62,340 ≈ 47%`
divides a MAIN-seat /context numerator by the SUBAGENT-seat by-difference denominator — exactly
the cross-seat, cross-channel mix this same commit forbids at rows 5c/5e and that §8.5 declares
irreconcilable. Restated within one channel: 17,363 × (4/2.6187) = 26,522 = 42.5% of the
62,340-tok seat, or ~26.5% against ADR-3's own ~100k denominator. Both readings put the
repo-owned share BELOW the 29-39% band, not above it, so the verdict is now "GAP-FOUND —
measured low, consistently across the conversion change" instead of "direction unresolved".

MINORs:
- §8.1 gave "two reasons, both binding" for not reconverting §2; one was FALSE. Concatenating
  the five files the ratio was measured on gives 69,453 B and row 1's published 17,363 est-tok
  is exactly 69,452 B / 4 — the same population, byte for byte. The claim is withdrawn in place
  and the surviving reason (re-derivation is beyond an addendum) is named as the only one. An
  unverified escape clause is a stronger shield than the correction it blocks, and this one was
  steering DECISION-NEEDED #4.
- DECISION-NEEDED #5 Option A's "wrong by roughly 2.4x" over-extended: 2.41x is the main-seat
  ABSOLUTE; the share moves 77.8% -> 49.5%, i.e. 1.57x, and the subagent-seat 68.4% is untouched
  because /context cannot run inside a subagent.
- §8.2 called the whole 29.4k memory block repo-owned; 2,764 of it is host-side (~/.claude
  CLAUDE.md 964 + MEMORY.md 1,800 = §2 rows 2 and 3). Repo-owned is 26,700 = 45.9% of the head.
- The §4 sweep had reached R1 and R5 but not R4, whose premise the paste contradicts: R4 rests
  on the harness truncating the skills listing "to a ~2k budget", while /context measures the
  injected block at 8.9k — essentially the un-truncated source-side ~9.1k. Surfaced for S-I,
  not re-derived here. R4's "129 SKILL.md files" also carries no reproducing command and a
  recount gives 112, so the population is marked UNVERIFIED.

Structural: absorbing the above pushed the patch to 602 lines, over the repo's 600-line markdown
gate. Trimming to 599 would be gaming the gate, so §8 is split into a companion patch,
2026-08-07-s-h-p14-context-addendum.md, with §8.x numbering preserved so every cross-reference
already written stays valid. Parent 435 lines, addendum 196.

Prior-art: skipped — review-absorption edit plus a size-driven split of an existing research patch, no new capability
…e 13)

The pre-push principle-13 gate correctly rejected the new patch: a research patch must carry an
actual §1.7 self-review, not merely name the section. Added Forward + Backward + T15.

The backward-check is a real outward sweep, not a restatement of this diff — the change class is
"a post-merge artefact that revises figures already published in a merged research patch", and
six surfaces are verdicted, of which four are GAP-FOUND and left unedited by ownership:
- the sibling p3d-p11 patch shares the falsified 4 B/t constant at its §5/§7/§8, so §8.1 applies
  to it identically — named in DECISION-NEEDED #4's Option A as required scope;
- the token-economy spec's tag convention (the constant under one of its tags is wrong);
- ADR-3 (repo-owned share measures below its 29-39% band under BOTH conversions);
- the S-A profile patch (closed historical artefact, its authoring session owns it).

T15 records the reflexive fact that this file exists only because the parent hit the 600-line
markdown gate — a document about document cost split by a size discipline.

Prior-art: skipped — self-review section required by principle 13 on an existing patch, no new capability
… round's own replacement figures

The split is sound (parent 463, addendum 261, all 43 §8.x cross-references resolve) and the
addendum's §1.7 backward-check verified as a real outward sweep. But three of round 3's four
replacement figures were themselves defective, plus a new challenge block that reversed a
downstream premise on an invalid comparison.

MAJOR — "Both readings put the repo-owned share BELOW the 29-39% band" is arithmetically false:
42.5% > 39%. And 42.5% is a share of the 62,340-tok SUBAGENT seat while ADR-3's band is
denominated on ~100k, so it is not band-comparable at all. Round 3 replaced a cross-SEAT mix
with a cross-DENOMINATOR one. Now stated from the directly measured figure with both traps
recorded inline so it is not re-derived wrongly a third time.

MAJOR — 26,522 was derived by applying §8.1's SEVEN-file aggregate ratio (2.6187, inflated by the
one host-side Russian-text outlier at 3.32 B/t) to row 1's FIVE-file population, while the
addendum measures that exact population directly at 26,700 (five-file ratio 2.6012). One commit,
two values for one block. The measured figure now supersedes the derivation: 26,700 = 26.7%
against ~100k (band-comparable, below the band) and 42.8% of the subagent seat (not comparable).

MAJOR — round 3's "29.4k is not repo-owned" fix was applied in §8.2 but not swept: §8.4 (the
S-D′-facing ranking section) and §4 R5's REVERSED note both still read "29.4k, 50.5% repo-owned",
overstating the own-able block by 2,764 tok at the one site a downstream stage reads. Both fixed
to 26,700 = 45.9%. Third site of the same class: "six ASCII-dominant repo files" counted
host-side MEMORY.md as a repo file.

MAJOR — the R4 CHALLENGED block concluded the skills listing "appears not to be truncated at
all", comparing the /context-measured 8.9k against a ~9.1k figure that is a 4 B/t estimate this
same commit declares low by 1.53x. In one constant: 41,057 B / 2.6187 = 15,678 tok, so 8.9k is
~57% of source; independently the snapshot lists 74 entries against a 112-file population, ~66%.
Both channels say REDUCED. The supported half survives — the ~2k budget premise is wrong by ~4x —
and that, not "no truncation", is what is routed to S-I.

MINORs: §4 R2's "until then / which R1 would settle" was stale once R1 discharged (now PARTLY
SETTLED, with the evidence stated as non-conclusive and the row keeping its UNMEASURED pricing
rather than gaining a "0"); §7's S-I-kickoff backward-check verdict was not re-adjudicated
although this commit moves that kickoff's premise (now GAP-FOUND, routed not edited); the T3
demand for a reproducing command was applied to the 129 being corrected but not to the 112
correcting it (command now published beside it).

The addendum's §1.7 now records the method failure rather than only the rows: four rounds, four
sweeps driven by the last review's list, each re-failing on whatever the list omitted — T21 in
its own-work form.

Prior-art: skipped — review-absorption edit on existing research patches, no new capability
…them a fourth time

Round 5 found 2 MAJOR, both again in the previous round's replacement figures. That is four
consecutive rounds where a hand-revised quantitative claim was itself defective, so this round
changes method: the unsupportable claims are withdrawn rather than corrected again.

MAJOR — the ADR-3 verdict was denominator-SELECTED, not measured. 26,700 has four defensible
denominators and they disagree in direction: 29.99% of this seat's own 89,019 first-turn total
(INSIDE the 29-39% band), 26.6% of the 60-session median (below), 45.9% of the /context resident
head (above), 42.8% of the subagent seat (above). Rounds 3-5 each picked one and each pick was
defective — cross-seat, then cross-denominator, then ratio-transferred-across-populations. The
verdict is now WITHDRAWN with all four denominators tabled and no verdict issued, and the choice
raised as DECISION-NEEDED #6 (which cannot be settled independently of #5, since the options
differ precisely by the ~30.8k dispatch-prompt gap #5 records).

MAJOR — the "74 listed entries / 112 files = 66%" corroborating channel is WITHDRAWN entirely.
The numerator is provably not a subset of the denominator: the two largest listed entries in the
capture, dataviz (~380) and claude-api (~360), have no SKILL.md anywhere, as do >=14 other
built-ins. The denominator is an unfiltered find carrying marketplace/cache duplicates, vendored
node_modules files, worktree copies, packages/core fixtures and uninstalled catalogue rows. A
ratio across two different sets measures nothing; publishing it would be the estimate-dressed-
as-measurement T-SH-A forbids.

MINORs: the "~57% of source" precision is withdrawn to direction-only — it swings 56% to 87%
across the four conversion constants in play, and the SKILL.md corpus is itself multi-byte-heavy
(six skills carry Russian descriptions), so no constant is defensible for it without measuring
that corpus. The 112 recount is no longer offered as a correction: publishing the command is
necessary but not sufficient, since the command must already exclude what the claim is not about.
measure-always-on.sh's "21-28%" gained the re-adjudication marker every sibling surface had.

The §1.7 note previously NAMED T21 while committing it. It now states plainly that this round's
sweep was list-driven too, that its hunks map one-to-one onto round 5's findings, and that the
class-driven counter T21 prescribes is what the five cold audit rounds have been doing while the
author-side sweep never became class-driven. It also records the second method finding: three
attempts to repair one comparison failed because the comparison had four denominators, and the
correct response was withdrawal.

Class sweep applied to the withdrawal itself: every site carrying a listing share was found by
grep and corrected, not only the one the audit named — the §7 S-I re-adjudication repeated the
withdrawn 57%/66% pair and now reads direction-only.

Prior-art: skipped — review-absorption edit on existing research patches, no new capability
…ate ranges, never magnitude words

Round 6 confirmed the round-5 withdrawal is complete (57%/66% survive nowhere; 74/112 only inside
their own WITHDRAWN notices; no third site) and re-derived all four tabled shares as correct. It
then found three MAJORs, all again in this round's own replacement wording.

MAJOR — withdrawing "~57% of source" to "a minority of source" INVERTED the claim. Under every
constant the same note lists, the injected 8.9k is 56.4% / 56.8% / 72.0% / 86.7% of source — a
majority — and 51.4% against the pre-S-I byte count. A magnitude word is not a weaker form of a
number, it is a different claim. Both sites now carry the explicit range and NO magnitude word;
the withdrawal rule is stated so the next editor does not substitute another adjective.

MAJOR — the new measure-always-on.sh re-adjudication claimed the measured 26,700 supersedes the
"21-28%" pair. Wrong on the NUMERATOR, not the denominator: 26,700 is the pre-S-G five-file set
(pinned byte-for-byte in §8.1) while the "~21%" member is the post-S-G set. No denominator choice
repairs a numerator mismatch, so no restatement is offered at all — the bound is unverified here
and both the surface and the post-S-G measurement stay S-E's.

MAJOR — "Options A/B and C differ precisely by the ~30.8k gap" holds only for A (89,019 − 58,200
= 30,819). B differs by 42,329 and is a 60-session median set against a gap measured on one
session, so B compounds #5 with a population change rather than restating it. Corrected in place.

MINORs: the "six skills carry Russian descriptions" clause is DROPPED rather than corrected — two
greps disagreed (6 vs a repo count polluted by node_modules), and the sentence two lines above
faults another figure for lacking a reproducing command, so publishing an unverifiable one there
was the same defect. "#6 below" pointed above. DECISION-NEEDED #6 is now propagated to every
enumeration that had stopped at #5: the §2 headline warning, R5's REVERSED note (which quotes
45.9% — one of #6's four tabled options, now labelled as such), the §8 pointer, the addendum
header and its §1.7 obligation count.

The §1.7 note also records that this round's two records disagreed about whether the sweep found
an unnamed site: the commit message was right, the paragraph was wrong. The S-I re-adjudication
was found by the author's own class grep. Honest summary now stated: list-driven for five rounds,
class-driven for exactly one item.

Prior-art: skipped — review-absorption edit on existing research patches, no new capability
…ad of restating it an eighth time

Seven cold rounds, seven REVISEs, and rounds 4-7 each found the MAJOR in the PREVIOUS round's own
replacement wording. The class never changed: deriving a quantity across mismatched populations,
denominators or conversion constants. This round stops deriving rather than deriving better.

MAJOR (round 7) — the injected-vs-source share published last round has exactly the defect the
adjacent paragraph withdraws another channel for: its numerator is the harness total for 74
LISTED entries, its denominator a byte sum over 129 SKILL.md FILES, and the same note proves
those populations differ (dataviz ~380 and claude-api ~360 are in the numerator and have no
SKILL.md at all). All four attempts at that share — a 66% population ratio, a ~57%, a "minority",
a 56.4-86.7% range — are now withdrawn together with one stated reason, and NO figure, range or
magnitude word replaces them. What survives is the comparison that needs no denominator: 8.9k
measured against the spec's ~2k budget.

MAJOR (round 7) — this file established last round that 26,700 is the PRE-S-G five-file set, then
used that fact in exactly one place while three others kept building on it: §8.2 reading 1 and
§8.4 ranked S-D′'s "top-ranked own-able target" from it, and R5 REVERSED overturned a
recommendation with it. The dominant file in that measurement, ai-laziness-traps.md (26,387 B), is
one S-G already replaced with a digest a quarter its size. So no ranking, no "top lever" and no
share of the current head is derived from the snapshot; what it establishes is the narrower and
still-useful claim that repo-owned memory files were the largest resident CLASS on the seat
measured. R5's reversal is downgraded to SUSPENDED: its premise moved, but its opposite is not
established either, and a post-S-G re-measurement is S-E's.

MAJOR (round 7) — §8.4 still said "#4 and #5", the site round 6 named live. Now #4/#5/#6, and §6's
Forks bullet — which still listed three — records the real inventory: six forks, five open.

MAJOR (round 7) — the §1.7 note asserted the §7 S-I site "was found by the author's own class
grep, not by any review naming it". False: round 6's watch-list names that exact line. The claim
is retracted in place, and the note now states plainly that the author-side sweep never became
class-driven in any round.

MINORs: "no magnitude word is used here" was false of its own sentence; the §8 pointer omitted
§8.6; "five rounds running" and "round 5's findings" were inherited prose that silently
re-pointed to a different commit.

Prior-art: skipped — review-absorption edit on existing research patches, no new capability
…asis says, repair the broken substitution

Round 8 recorded the first CLEAN on W-5, the cross-population-share criterion that produced the
MAJOR in rounds 4, 5, 6 and 7. The remaining findings are a different and narrower class.

MAJOR — §8.4 stated its withdrawal broader than the §8.2 reading it cites as basis: §8.2 withdraws
shares "of the current head", §8.4 restated that as "every share of it [the pre-S-G set] is
withdrawn" — while four surfaces still publish exactly such shares (§8.2's class table, §8.6's
four-denominator table and their two parent-side twins) and while holding #6 open, whose entire
subject is which share to publish. Narrowed to match: the pre-S-G shares stay valid AS pre-S-G
shares and are not withdrawn; what is withdrawn is any share of, or top lever for, the CURRENT
set — which is what a subtraction map acts on. #6 is now named as the question of which
denominator a pre-S-G share is quoted against.

MAJOR — propagating the share-withdrawal into the §7 S-I surface broke the sentence: "the listing
is nonetheless reduced to measured at 8.9k injected" left a dangling verb phrase, asserted 8.9k
twice, and kept the direction word "reduced" that R4 forbids six lines into its own text. Rewritten
to carry R4's own closing position: the budget premise is wrong by ~4x, and NO claim is made about
truncation either way.

MINORs: §1.7's marker inventory still read "R5 REVERSED" after this round renamed it SUSPENDED;
the -20,782 B set cut was attributed entirely to the traps->digest swap, which accounts for
-19,684 B (the rest is two other files in the same trim); §8.4 called §8.1's measured B/token
aggregate an "identity" alongside §8.2's exact arithmetic one, upgrading a 2.37-3.32 empirical
average to an exact relation in the round whose purpose was the opposite; "the file that dominates
this measurement" is 9.8k against CLAUDE.md's 9.3k, so it is the largest single file, not a
dominant one.

Prior-art: skipped — review-absorption edit on existing research patches, no new capability
…o §1.7, drop the wrong-side share label

Round 9 resolved five of round 8's six findings and returned one MAJOR of the same W-17 class at a
site the previous commit did not reach, plus one MINOR in the wording it introduced.

MAJOR — §8.4 narrowed its withdrawal to "any share of the CURRENT set", but the §1.7 T15 paragraph
still carried the pre-round-8 breadth: "(§8.2 reading 1, whose share figures are withdrawn as
pre-S-G)". The file therefore issued two incompatible instructions about the same table to the same
consumer, and the §1.7 form also dropped the "of the current head" qualifier its cited basis carries
(§8.2 reading 1). Restated to match that basis exactly: a pre-S-G measurement from which no share of
the current head is derived. Enumerated every withdrawal statement across all three S-H patches
(grep -n withdraw → 6 hits: addendum :69, :159, :290, :327; parent :262, :297); this was the sole
over-broad survivor — :262 withdraws the injected-vs-source share, :297 withdraws the R5 reversal,
both correctly scoped.

MINOR — the exemption introduced last round read "remain valid as shares of that pre-S-G set". That
is exact for §8.2, whose denominator IS the pre-S-G resident head (58.2k), but inverted for §8.6,
where the pre-S-G block (26,700) is the NUMERATOR and the four denominators are seat totals — the
relation the same paragraph states correctly two lines later. Replaced with a form true of both:
computed from that pre-S-G measurement, each against the denominator its own table names.

Both edits are subtractive/narrowing and introduce no figure, ratio or magnitude word — the
strategy that first produced a CLEAN on W-5 at round 7.

Prior-art: skipped — review-absorption edit on existing research patches, no new capability
@artyhoo
artyhoo merged commit 447643a into staging Aug 7, 2026
39 checks passed
artyhoo added a commit that referenced this pull request Aug 7, 2026
…rable gate MET (#1267)

Rev 7 left five sites asserting that S-L is unmerged, the loudest being
`…-s-d-prime/kickoff.md:361` «The stage still cannot start, for one remaining
reason: **S-L is not merged**». S-L merged as PR #1263 on 2026-08-07T12:50Z, so
an executor reading the kickoff in full — which §0 requires — hits a stop-text
that is now false. Dispatching against it would be `#dispatch-before-staging`
in the other direction: the input on staging says «do not start».

Retired at all five sites, three in the stage kickoff (`:1` header, `:310`
§5 consequence line, `:361` dispatch status) and two in the umbrella (`:97`
stage-table cell, `:362` S-L ordering paragraph).

Scope: no deliverable, no permitted-file set, and no acceptance criterion
changes. The rev-8 dispatch-status paragraph additionally carries forward the
one binding thing S-L's §5 says about this stage — «A re-ranking is not a
rescale — S-D′ must re-derive rather than multiply through»
(`docs/meta-factory/research-patches/2026-08-07-s-l-recalculation.md:509-510`),
plus the note that both hook injects are levers a `/context`-ordered list ranks
at zero (`:499`) — because an executor that multiplies through a uniform factor
preserves order by construction and would hide exactly the effect S-L found.

Both gates verified mechanically, not from the umbrella prose (the recurrence
that memory `verify-before-claim-family` records):
  S-E  #1237 MERGED 09:39Z; meter present, scripts/measure-always-on.sh:10-11
  S-H  #1239 MERGED 00:06Z + #1249; P11 returned a real absence with a
       discriminating control, NOT INCONCLUSIVE (…-p3d-p11.md:431,440)
  S-L  #1263 MERGED 12:50Z; §5 «Spec reach — what this does to the S-D′
       ranking» (…-s-l-recalculation.md:493-510)

Umbrella kickoff stays at 592 lines (600-line pre-commit gate): the two edits
are in-place rewrites, no appends.

Prior-art: skipped — doc-only correction to two dispatch-input kickoffs; retires a stale gate claim, adds no capability, no dependency, no code.

Co-authored-by: Test <test@example.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant