Skip to content

chore(skills+agents): cleanup batch from 2026-05-16 skills+agents audit - #64

Merged
artyhoo merged 4 commits into
mainfrom
chore/skills-agents-cleanup-wave-2-2026-05-16
May 16, 2026
Merged

chore(skills+agents): cleanup batch from 2026-05-16 skills+agents audit#64
artyhoo merged 4 commits into
mainfrom
chore/skills-agents-cleanup-wave-2-2026-05-16

Conversation

@artyhoo

@artyhoo artyhoo commented May 16, 2026

Copy link
Copy Markdown
Owner

Summary

Wave 2 cleanup batch executing D-AuditC-1 R-1, R-3, R-4 (verify-only), R-6, R-7 + D6 + D-AuditC-6 from the 2026-05-16 strategic D-items dialogue (decisions.md, gitignored). All edits are documentation / metadata / cleanup; no rule introduction.

  • C1fix(agents): remove surviving broken self-testing-docs echo (acfdad8) — symmetric to c8fc153; line-41 echo referenced absent agents/references/ path.
  • C2docs(agents+CLAUDE): consumer-facing context note + ownership row per D-AuditC-6 (0a1e3ad) — clarifies docs-auditor.md (scripts/audit-ai-docs.sh) and best-practices-sidecar.md (.ai-factory/RULES.md) source-vs-consumer gap; adds CLAUDE.md Artifact Ownership Contract row.
  • C3docs(skills): mark tool-bootstrapping Wave 5.3 as landed (c0efa1a) — tense correction (landslanded).
  • C4docs(research-patch): §17 think-time-gate erratum cross-link per D6 (087bc7c) — preserves maintainer-flagged residual doubt (H2 vs H10 re-eval deferred to Phase 11+).

Companion user-home edits (NOT in this PR): R-2 was a no-op (pre-cleaned), D-AuditC-3 mv executed manually by orchestrator after sandbox-blocked subagent attempt — see Test plan / ATTN.

Test plan

  • Pre-push hooks pass cleanly (no --no-verify)
    • audit-ai-docs.test.sh: 9 pass / 0 fail
    • pre-push.test.sh: substance + Prior-art + §1.7 checks all PASS (16 pass / 0 fail)
    • hook-stub-completeness.test.sh: ✅
    • render-rules.ts --check: rules-table region up-to-date
    • test:principles: 10 files / 56 tests pass (714ms)
  • Verify probes captured for every edit (probe-before + probe-after with real grep/sed output)
  • No drive-by edits outside the named files
  • No rule-introducing change (§1.7 Skipped marker rationale below)

ATTN (deferred — not actioned in this PR)

  • deps-hash-check.sh is not registered in .claude/settings.json UserPromptSubmit hook (D-AuditC-1 R-3 observation). Scope-creep — surface for separate session.
  • H1 (R-2): ~/.claude/skills/ai-docs/SKILL.md had no token-economy references at execution time (pre-cleaned state, 0 grep matches). No-op.
  • H2 (D-AuditC-3): mv ~/.claude/skills/git-user-info-ui-design.md~/.claude/skills/projects/github-user-analytics/ — subagent sandbox blocked the move; orchestrator executed manually after explicit maintainer confirmation.

§1.7 Skipped: cleanup-only batch — no new rule, no new principle test, no new gate; CLAUDE.md row documents existing rule (D-AuditC-6), not a new one. Forward+Backward scope (per phase-research-coverage.md §1.7) protects rule-introduction drift, not mechanical maintenance of existing artefacts.

🤖 Generated with Claude Code (orchestrator session, Wave 2 of 2026-05-16 D-items execution)

artyhoo added 4 commits May 16, 2026 23:25
The Step-2 graceful-degradation block in agents/docs-auditor.md echoed
'See references/self-testing-docs.md for the pattern.' — referencing a
path that does not exist (agents/references/ absent in source + consumer).
Symmetric to former line-12 instance fixed in c8fc153.

Surfaced by 2026-05-16 Wave 2 audit (D-AuditC-1 R-1).
… D-AuditC-6

Two consumer-facing agents reference paths populated only by the AIF
installer in consumer projects (scripts/audit-ai-docs.sh for docs-auditor;
.ai-factory/RULES.md for best-practices-sidecar). In the source repo these
paths are absent; the agents handle this via graceful degradation. Note the
expectation explicitly in each agent's description, and add a CLAUDE.md
Artifact Ownership Contract row clarifying ownership.

Per decisions.md D-AuditC-6 (2026-05-16). Wave 2 R-6 + R-7 + ownership row.
Wave 5.3 hook implementation has shipped. Tense correction only.

Per decisions.md D-AuditC-1 R-3 (2026-05-16).
Add a See-also erratum pointing to 2026-05-16-think-time-s17-gate-correction.md.
H2 vs H10 architectural re-evaluation with corrected Stop semantics is
deferred to implementation moment (Phase 11+); maintainer flagged residual
doubt — flag deliberately preserved, not silently absorbed.

Per decisions.md D6 (2026-05-16).
@artyhoo
artyhoo merged commit 7c19477 into main May 16, 2026
37 of 38 checks passed
artyhoo added a commit that referenced this pull request May 22, 2026
…n sequencing plan (#160)

Records the autonomous Queue-mode N8 R-phase outcome into the sequencing plan §0 snapshot:
- §0 N8 row: 🟡 «plan committed; impl pending» → ✅ R-phase DONE (#158→staging, Worker→Reviewer GO→anti-collusion), A-phase 🔲 pending D1/D2/D3
- §0 «what remains»: N8 A-phase gated on maintainer D1 (cost-lever A/B/C) / D2 (local-model bench) / D3 (offload priority), surfaced in findings §7; SSOT #64#68 ship with capability
- §2 Track-1 1.1 + §4 first-launch: marked DONE (recommendation actioned)

No status invented — reconciled against PR #158 + the n8-rphase-findings deliverable.
artyhoo added a commit that referenced this pull request May 22, 2026
#165)

* docs(meta-factory): N8 A-phase gate — list D1/D2/D3 explicitly in Track-1 row + dep diagram (§7 link)

Track-1 row 1.3 and the dependency diagram showed only Depends-on 1.1;
the D1/D2/D3 maintainer gate lived only in the §0 prose snapshot. An
orchestrator picking work from the operational table/diagram could start
A-phase impl before the cost-lever / bench / priority decisions are made.
Now both operational touchpoints name D1·D2·D3 + link findings §7.

Prior-art: skipped — doc-only edit, no new capability (sequencing-plan prose, no code/dep/artifact added).

* docs(meta-factory): clarify N5/N6b readiness + SSOT-ID collision in §0 (anti-confusion)

Sessions kept reading 🔲 as 'blocked' regardless of cause. Disambiguate §0:
- N5: 🔲 **BLOCKED** — gated on N7 live-dogfood trial (decision=C done, repo-side applied, live trial pending) + N2; can't give back before dogfooding reveals what's unique.
- N6b: 🔲 **NOT blocked** — deps met (N3 portable TS-core = packages/core/hooks/pre-push.ts+checks/; the 7 .claude/hooks/*.sh are thin CC-event glue, not engine; + N6a ✅). Sequenced 'last' by choice, startable anytime.
- 🔲 legend added: blocked (unmet dep) vs deprioritized (deps met, startable) — the exact distinction sessions confuse.
- ⚠ SSOT-ID collision flagged: N7 took #64/#65; N8 findings §7 proposed #64#68 → N8 A-phase renumbers to #66+ at admission.

N7 row left untouched (freshly applied by its owning session). Status reconciled against origin/staging + Wave-10 TS file inventory.
@artyhoo
artyhoo deleted the chore/skills-agents-cleanup-wave-2-2026-05-16 branch May 22, 2026 18:09
artyhoo added a commit that referenced this pull request May 22, 2026
…E Superpowers + SSOT #64/#65 + retention=A coexist (#166)

N7 (DECISION=C) repo-side application. Source-verified the Superpowers
adoption against shipped SKILL.md (dual-channel DeepWiki + raw WebFetch,
2026-05-22) instead of assuming from name.

- .claude/rules/parallel-subwave-isolation.md §4: drop the homegrown
  AST-detection build-target; REFERENCE Superpowers `using-git-worktrees`
  (Step 0 GIT_DIR!=GIT_COMMON skip → compatible with isolation:"worktree").
  Stays Class C; new §5 §1.7 forward/backward note.
- prior-art-evaluations.md: append SSOT #64 (SDD, ADOPT process-layer) +
  #65 (using-git-worktrees, ADOPT + REFERENCE). NOT #60/#61 — those were
  taken by the parallel channel-selection wave (the exact ID-collision §6
  warned about). "#27 = Git-Isolation" was a misattribution (#27 is
  HANDOFF_MODE) — corrected, void #27-update skipped.
- N7 patch §9 closure: source-verification findings, retention verdict
  A (coexist), ID corrections, classifier-blocked items.
- roadmap banner reconciled (DECISION=C set, original framing preserved
  append-only); sequencing-plan N7 status updated.

Retention=A: SDD review is post-implementation, orchestrator Phase -1 is
pre-dispatch + owns quota/Mode/bootstrap meta-layer SDD lacks → coexist.

Blocked (maintainer-applied): global Superpowers install + orchestrator-
skill annotation — both denied by the auto-mode classifier (untrusted
external code / self-modification). Live-dogfood trial pending those.

§1.7: Forward — demotion complies with build-first-reuse-default.md (REFERENCE-over-BUILD) + no-paid-llm-in-ci.md (using-git-worktrees pure-git); Backward — scope-reducing, only edited bullet is parallel-subwave-isolation.md:37 (§4), SSOT cross-ref at prior-art-evaluations.md row #65, full self-reflexive note at parallel-subwave-isolation.md:46
Prior-art: skipped — N7 process-layer dogfood adoption; SSOT #64/#65 appended documenting ADOPT verdicts, no new dependency or capability code added
artyhoo added a commit that referenced this pull request Jun 3, 2026
… (Stage 1) (#403)

Adds .claude/skills/dispatcher/SKILL.md (~195 LOC) wiring 4 existing CLI
primitives (dispatch.ts / harvest.ts / questions.ts / answer.ts) into a
dispatch→monitor→Q&A→harvest→Phase-1→stage-gate→advance loop. DN-A = dual
(CC-native primary + portable fallback via capability-check, not brand-name;
@dual-pair: dispatcher-skill). DN-B = Option A behavior: technical forks
resolved autonomously via superpowers:brainstorming on CC; strategic forks
surfaced to operator via questions.ts. Discrimination discipline baked into
prose (ask-question-reminder.sh is operator-internal, not shipped).

Prior-art: prior-art-evaluations.md#111 (BUILD — no upstream covers cross-boundary dispatch→monitor→Q&A→harvest→stage-gate loop; SSOT consult + DeepWiki ≥3 + WebSearch ≥3 in rphase §c; T16 verified vs #64/#67/#88).

§1.7: forward-check applied — BFR=BUILD via all 5 layers (SSOT consult + DeepWiki≥3 + WebSearch≥3, dispatcher-skill-rphase.md:223), dual not cc-only per dual-implementation-discipline.md:3, zero-paid-LLM per no-paid-llm-in-ci.md:1, T16 problem-class X-vs-Y at SKILL.md:189; backward-check sweep — complementary to pipeline/SKILL.md:313 (no supersession), consumes runtime-bridge primitives + park.ts #109, SSOT #111 appended at prior-art-evaluations.md:181.
artyhoo added a commit that referenced this pull request Jul 2, 2026
…-development (#859)

BFR self-correction (operator-flagged #parallel-evolution-creep): the first night-mode re-described SDD's coordinator + implementer + dual-reviewer loop (~70% of SSOT #64). Slimmed from ~160 to ~44 lines — the loop is now DELEGATED to superpowers:subagent-driven-development, and the skill owns ONLY the overnight delta SDD lacks: unattended autonomy/fork policy, quota/backoff resilience (ADAPT AIF watchdogs #45), Workflow context-economy, verification discipline, the verified diff-visibility harness fact, and the unsupervised terminal condition. Header + paired-negative retained; principles 09/14/15 green.

Prior-art: prior-art-evaluations.md#64 (Superpowers subagent-driven-development, ADOPT — night-mode now explicitly a thin overnight adapter over the SDD loop, not a re-implementation) + prior-art-evaluations.md#45 (AIF watchdogs self-healing, ADAPT — the quota-backoff resilience basis).

Co-authored-by: t <t@t.co>
artyhoo added a commit that referenced this pull request Jul 3, 2026
…capability-reuse-auditor (#863)

Ships **source-before-shape** — an edit-time discipline catching two recurring AI-laziness failures at authoring time, closing a recursive-self-application gap in the project's own operating rules.

Origin (2026-07-02, operator-confirmed cross-session recurrence = promotion trigger):
1. BFR-reinvention — PR #858 shipped a night-mode SKILL.md re-describing the loop SSOT #64 owns; its trailer said "ADAPT #64" while the body re-described it, and it passed principle 11 F1 (which checks trailer presence, not reuse substance). #consult-as-trailer-not-input.
2. scope-from-memory — a launch prompt scoped from recall, not the spec (D1 → B). #claim-from-memory-not-source.

Mechanism (judgment → injection, not a gate):
- Layer A: .claude/rules/source-before-shape.md carries globs/inject markers → the existing inject-matching-rule.sh surfaces the reminder at edit-time (.claude/skills/**, agents/**, .claude/orchestrator-prompts/**). REUSE, zero new engine. Honestly disclosed once-per-session limitation (best-effort first-touch nudge).
- Layer B: agents/capability-reuse-auditor.md — AI-agnostic overlap + trailer↔body auditor (no-paid-llm), doing the semantic pass F1 cannot.

SSOT #196 (ADAPT, records the 6-item BFR-consult); principle 09 REQUIRED_HEADER_DOCS +2; install.sh SHIPPED_DOCS +1 + regenerated byte-identical baselines; AGENTS.md rule-index +1; self-reflection research-patch.

Independently reviewed (3-lens Workflow: 2 approve + 1 adversarial revise applied). Verified: principles green, byte-identical 8/8, injection dogfooded live. Candidate T-trap #consult-as-trailer-not-input surfaced-not-applied to ai-laziness-traps.md (dedicated rule home instead).

§1.7: Forward+backward in source-before-shape.md §6 — complies with no-paid-llm-in-ci.md §1, build-first-reuse-default.md §4, phase-research-coverage.md §1.11, doc-authority-hierarchy.md §2-§3; registered at packages/core/principles/09-doc-authority-hierarchy.ts:54; self-applies (SSOT #196 consult drove the shape, not memory).

Prior-art: prior-art-evaluations.md#196 (source-before-shape mechanism — ADAPT: REUSE inject-matching-rule.sh channel + BUILD the AI-agnostic capability-reuse-auditor; dedup-first ADOPT-VOCABULARY; code-clone tools REJECT on T16 problem-class miss; Superpowers writing-skills REFERENCE).
artyhoo added a commit that referenced this pull request Jul 18, 2026
…1032)

* feat(skill): claude-glm-executor-handoff + advisor-consult sub-form

Ships a thin-adapter skill for the Claude→GLM cross-model dispatch edge
inside aif-handoff, plus an advisor-consult sub-form on the existing
REPORT BLOCKER field, plus a research-patch provenance.

- new .claude/skills/claude-glm-executor-handoff/SKILL.md (130 lines):
  5 verified GLM-5.2 facts (D1-D5, source-grounded to Z.ai docs) +
  3 refuted-folklore entries + 6-block input contract + status
  translation + recovery protocol + §5 honest-gaps marker (5 designed-
  not-proven claims). Paired-negative block per principle 15.
  Subordinates to night-mode/SDD/orchestrator-worker-discipline.

- new docs/meta-factory/research-patches/2026-07-18-claude-glm-
  executor-handoff-facts.md (89 lines): provenance for the skill.
  Verified-facts table + refuted-folklore log + probe design (P1/P2/P3).
  Substrate (P4) resolved same-day: aif-handoff supports per-agent
  model: frontmatter natively.

- agents/orchestrator-worker-discipline.md (+50 lines): new Advisor-
  consult protocol section. Adds BLOCKER: advisor-consult: prefix
  sub-form on the existing REPORT schema (no schema extension).
  Coordinator routing table for 3 cases. Caps at 2 per task.
  Subordinates to night-mode delta 7 for overnight cases.

- .claude/skills/night-mode/SKILL.md (+2 lines): one-line pointer
  to the new skill, anchored on the existing Harness portability
  paragraph.

Prior-art: SSOT #64 (SDD, ADOPT — the loop), SSOT #201 (Anthropic
Advisor tool, ADAPT — the strategy generalised from night-mode delta 7
for non-overnight dispatch), SSOT #200 (harness-config emission,
referenced in skill D6).

§1.7: forward+backward applied — agents/orchestrator-worker-discipline.md:51 introduces Advisor-consult protocol (new discipline sub-form); .claude/skills/claude-glm-executor-handoff/SKILL.md:99 carries the §5 honest-gaps marker gating contract-section promotion; backward sweep of .claude/skills/**/SKILL.md and agents/*.md confirmed no re-description of night-mode/SDD/orchestrator-worker-discipline surfaces (Layer B capability-reuse-auditor verdict THIN-ADAPT).

Honest gaps: skill §5 lists 5 designed-not-proven claims; advisor-consult
reliability is similarly under trial. Promotion to hardened contract
requires live probes recorded as follow-up research patches.

Two-round review: Reviewer A (concept, bottom-up) + Reviewer B (execution, top-down) on draft v2, then cold /reviewer Mode 3 pass on final state. All 4 BLOCKERs from Reviewer B applied (status enum DONE|BLOCKED|PARTIAL; pay-as-you-go endpoint /api/paas/v4/; cached pricing cut; 5 smuggled-unverified claims sourced). Layer B capability-reuse-auditor verdict: THIN-ADAPT (body genuinely subordinates, no re-description).

* chore(synth-bundle): rebuild after skill addition

Pre-push hook caught drift — claude-glm-executor-handoff/SKILL.md
addition changed the synth-and-wire bundle output. Rebuild to match.

* chore(backends): npm capability-matrix eslint freshness — 9.39.4

Pre-push freshness gate caught drift: matrix claimed eslint 10.4.0,
locally-resolving eslint is 9.39.4. Live firing test passes on 9.39.4
(5/5) — diagnostic format stable across the version range.
Unrelated to the GLM handoff PR; bundled to unblock the push.

* revert(backends): npm capability-matrix back to eslint 10.4.0 (CI truth)

Previous commit 8c3e636 set the matrix to 9.39.4 to match locally-
installed eslint, but CI does a fresh install and resolves 10.4.0.
Pre-push freshness gate passed locally but failed in CI.

Root cause: local node_modules was stale; the matrix was correct as
originally committed (10.4.0 matches what CI sees). Fix: bring local
eslint in sync (npm install eslint@10.4.0) AND keep the matrix at
10.4.0.

* chore(synth-bundle): rebuild after npm install eslint@10.4.0

---------

Co-authored-by: t <t@t.co>
artyhoo added a commit that referenced this pull request Jul 30, 2026
…tpocock/skills + superpowers (#1183)

Answers the operator's «кажется мы опять переизобрели велосипед» for the
orchestration contour (/arch, /pipeline, /dispatcher, night-mode, aif-doctor,
packages/runtime-bridge), closing the explicit gap PR #1181's raw material
left open (it read DeepWiki summaries, never the shipped SKILL.md bodies).

Answer: no for the substrate — the runtime is lee-to/aif-handoff
(AifHandoffBackend.ts:1-2 "adapter for the lee-to/aif-handoff runtime"; 0 of
2175 TS lines implement a runtime), the executor loop is superpowers SDD
(SSOT #64), the ideation loop is superpowers brainstorming, and
disable-model-invocation is convergent with mattpocock's primary mechanism.
Yes, narrowly, for two claims in our own bodies:

- arch/SKILL.md:79 asserts "no design-review skill exists" upstream, but
  skills/brainstorming/spec-document-reviewer-prompt.md ships in 6.1.1 AND
  6.2.0 (orphaned, referenced by no skill body), and its loop was retired at
  v5.0.6 after an A/B finding "identical quality scores" for ~25 min overhead.
- night-mode/SKILL.md:15,24 describe SDD's 5.1.0 two-reviewer roster (retired
  into one task-reviewer-prompt.md), overstate the whole-work-review gap
  (SDD:391-414 has a whole-branch Final Review), and silently cap rework at
  ~4 rounds against SDD:320's five.

Both are recorded, not fixed — §5 scopes them as proposals (F2/F3); skills and
rules are out of scope per the Artifact Ownership Contract.

10 per-capability verdicts on both BFR axes with a T16 problem-class line and
a falsifier each; three retain BUILD, one flips to WATCHLIST because SSOT
#111's search missed builderz-labs/mission-control (5.9k stars, created
2026-02-13 — before that evaluation).

Method: SSOT consult first (18 rows cited by ID), DeepWiki x4 across 2 repos,
WebSearch x3 phrasings, gh api + direct reads of 3 cached superpowers
versions. context7 excluded per BFR §3 tooling caveat.

Prior-art: prior-art-evaluations.md#230 (mattpocock/skills, verdict REFERENCE — per-capability KEEP NARROW/no-match, one already-convergent mechanism), added in this commit.
Prior-art: prior-art-evaluations.md#231 (superpowers retired spec-document-review loop, verdict REFERENCE — standing negative evidence against /arch §2), added in this commit.
Prior-art: prior-art-evaluations.md#232 (builderz-labs/mission-control, verdict WATCHLIST — fires SSOT #111's own revisit trigger), added in this commit.

Co-authored-by: Test <test@example.com>
artyhoo added a commit that referenced this pull request Jul 31, 2026
…mbrella + S-A kickoffs (#1189)

* docs(handoff): /arch v2 + context-pipeline session handoff — decisions, in-flight aif tasks, continuation protocol

Prior-art: skipped — session handoff document only, no new capability; decisions it records cite their own SSOT rows (#231, #207, #179, #64).

* docs(spec): /arch v2 + context-pipeline system design — layer model, pipeline arc, ADR-1..8

Step 5 of the 2026-07-31 handoff §4 protocol. Fable design authored on the
Opus research distillate (spot-checked, freshness-barred) and the Opus cold
critique (GO-WITH-PATCHES). All three critique blockers absorbed by
re-derivation: #231 over-read retracted (ADR-5), 5/5-K1 incident count
corrected to 2/5 and the primary/background split dropped (ADR-6), the
calibration falsifier given an oracle via shadow-A/B + pre-declared
threshold (ADR-5). M1-M7 absorbed as design constraints (population table,
bounded drill-down + distillate K-pass, K6 candidate/adjudicate split,
gate-channel re-route to pre-push/CI, operationalized bet falsifier,
L1/L2 boundary re-drawn, option spaces spanned).

Prior-art: skipped — design spec only, no new capability shipped

* docs(arch-v2): umbrella kickoff S-A..S-F + S-A stage-scoped dispatch input

Execution plan for the /arch v2 + context-pipeline track, derived from the
2026-07-31 design spec (ADR-1..8) by the Opus plan-writing seat.

Umbrella kickoff: stage table S-A..S-F with per-stage scope, dependencies,
tier classification (justified against CLAUDE.md's fixed criteria), acceptance
and implemented ADRs; dispatch protocol (4-arm in-flight probe, Phase -1 cold
review, bridge-profile marker rule with the fidelity-verdict precondition
quoted verbatim and re-verified at dispatch); calibration-ledger bootstrap
(ADR-5/6/8) with the ADR-8 token instrument named; cross-umbrella dependency
on token-audit S1 (S-E only, two gates: merged AND content-read).

The bottom seat + shadow-A/B station is marked active from S-B merge onward —
S-A predates the contract implementation and is covered only by Phase -1 plus
its own acceptance commands. Stated, not papered over.

Plan-writer objections (§4, per «who must write the plan cannot rubber-stamp
the design»): O-1 three wrapper drifts, not two, and one mis-described —
upstream ships brainstorming/spec-document-reviewer-prompt.md in 5.1.0/6.1.1/
6.2.0, night-mode:15's SDD roster does not match upstream, night-mode:29 cites
stale upstream line numbers; O-2 the skill-exists-by-name smoke catches none of
them and skips silently off-host; O-3 ADR-8's token metric had no named
instrument (aif task tokenTotal/costUsd, verified live); O-4 the ledger
principle test is vacuous before 5 rows; O-5 S-D's tier is a function of S-C's
verdict; O-6 the spec's marker condition drops CLAUDE.md's «produced by /arch».

S-A kickoff is stage-scoped and self-contained (handoff decision 11): W1-W6
with concrete file targets and a verification command each, host-verify
contract, descopes, §1.7 obligation in enumeration format, T-enumeration plus
three domain traps.

Prior-art: skipped — kickoff/plan docs only, no new capability

* docs(arch-v2): preserve track evidence artifacts — distillate, corrected idea, cold critique

The design spec cites these three as its evidence chain (distillate →
corrected idea → GO-WITH-PATCHES critique); they lived only in the
session scratchpad under /private/tmp, which does not survive a reboot.
Committed verbatim as session artifacts of the 2026-07-31 protocol run.

Prior-art: skipped — evidence-record docs only, no new capability

* docs(spec): point evidence-chain citations at the committed artifact files

Prior-art: skipped — link fix in a design doc, no new capability

* docs(spec): absorb token-audit S1 acceptance — fresh N2 numbers, ADR-3 falsifier fired

S1 (task c781e8a9, accepted 2026-07-31) measured the repo-owned always-on
set at 29-39% of the observed ~100k session-start total, firing ADR-3's
pre-registered falsifier: the budget gate's asserted quantity is re-scoped
to the repo-owned share (explicitly labelled), the harness remainder
routes to settings-recommendations, and the InstructionsLoaded
verification task doubles as the measurement-extension probe. N2 updated
to the fresher script-reproducible per-environment numbers (140,216 B
host vs 118,374 B container), replacing the older channel-level A7 pair.

Prior-art: skipped — design-doc correction on fresh measurement, no new capability

---------

Co-authored-by: Test <test@example.com>
artyhoo pushed a commit that referenced this pull request Jul 31, 2026
…n fix introduced

The adversarial re-review confirmed the previous round's fixes hold (permission
denial is now loud at all seven levels; all four symlink shapes are followed;
10 of 13 mutations killed) and found that the fix itself had introduced a new
false alarm, plus one test that did not test what it named.

MAJOR — a stray FILE in the plugins cache turned a healthy install into a loud
BROKEN, and worse, turned "nothing installed" into a FAIL. Probing
`<entry>/superpowers` without first probing `<entry>` reports ENOTDIR for any
plain file, and a single `.DS_Store` — one Finder visit — was enough. The root
cause was treating every non-ENOENT errno as an alarm, so the fix is a
distinction rather than another special case: an error now means "we were
DENIED the answer", never "the answer is no". EACCES/EPERM can conceal an
install and stay loud; ENOENT/ENOTDIR/ELOOP are definitive negatives and are
skipped. That resolves the same class at every level at once, including the
ELOOP false alarm on a self-symlink.

MAJOR — the symlink test named "marketplace AND version dirs" but symlinked
only the marketplace, which resolves through ordinary path resolution; a
faithful lstatSync mutation at the version and skills probes left it green. Now
one case per level, each also resolving a reference through the discovered root
so the root is proven usable rather than merely listed. The same mutation now
kills four tests.

MINOR — an unreadable skill directory under a readable root (mode 0444) blamed
the reference for a permission problem; the root is now named. A bare `return`
guard reported a green tick for a test that asserted nothing under root —
`it.skipIf` reports a skip. An unset HOME was a silent SKIPPED pass; "we cannot
work out where to look" is not "nothing is installed".

MAJOR (docs) — arch/SKILL.md's description and its §2 bullets assigned tiers to
the two review seats, while the seat-instantiation paragraph puts BOTH seats on
the mid tier whenever the authoring session is top tier — which §0 says it
always is. The advertised tiers could therefore never occur. Seats are now
altitudes only; tiers are stated once, where they are decided.

MINOR (docs) — "brainstorming dispatches an author-side spec-document reviewer"
overstated upstream: the prompt template ships, but nothing in the 6.2.0 flow
references it (grep over the installed skills returns no referencing line) and
that flow's spec pass is a self-review. Replacing a negative-existence overclaim
with a positive one would have repeated W4(a) in mirror image.

Surfaced, not fixed (outside the round-2 permitted file set): the same W4(b)
roster drift survives in pipeline/SKILL.md:320-323, which additionally credits
it to SSOT #64 vocabulary.

Verified on the host: 18/18 on the smoke, 71/71 with principles 09/12/14/15;
`.DS_Store` with and without an install both behave; the permission case stays
loud; the lstatSync mutation kills 4 tests; rule-index green; arch 122 lines,
description 1216 chars.

Prior-art: skipped — defect repair inside an existing test and one existing skill; no new capability, dependency, or engine introduced.
artyhoo added a commit that referenced this pull request Jul 31, 2026
…acceptance (#1192)

* feat(arch-v2-s-a): /arch v2 rewrite — research contour, membrane/K-pass, kill channels, W4 drift fixes, upstream-ref smoke

S-A stage of the arch-v2-context-pipeline umbrella. Implements W1-W6 per the
design spec §2 + ADR-4 + §4 item 1.

W1 — §1.5 research contour: trigger + explicit skip, research-spec template
(required pre-mortem + acceptance-criteria fields), execution with freshness
bar, distillation with GO/rework/kill verdict, seats as relative tiers.

W2 — membrane + K-pass + bounded drill-down (ADR-4): default consumption rule
with bounded recourse (not isolation), K1/K2 pass on each distillate before
consumption, drill-down capped at ≤3 per artifact with recording requirement.

W3 — cold definition promoted to a named single statement («did not author
AND did not receive authoring context»), referenced from §1.5 K-pass and §3
exit, not re-stated per section. Kill channels enumerated with cost ordering.

W4 — three wrapper drifts repaired at interface level per handoff decision 6
(no version pins, no line numbers, no upstream internals):
  (a) arch/SKILL.md:79 — negative-existence claim → upstream capability + our delta
  (b) night-mode/SKILL.md:15 — roster contents → capability delegated to SDD
  (c) night-mode/SKILL.md:29 — line-number citation → behaviour citation

W5 — packages/core/skills/upstream-skill-reference.test.ts smoke: asserts every
superpowers:<name> reference in .claude/skills/*/SKILL.md resolves to an
installed upstream skill directory. Honest scope — does NOT cover W4 drifts
(stated in test header). Environment-aware: discovers upstream by glob,
SKIPPED when absent. Paired-negative RED+GREEN+SKIPPED all exercised.

W6 — unique-filenames dispatch contract for parallel subagents added to §2
(handoff decision 13; near-clobber incident 2026-07-31).

Verification:
- grep §1.5|pre-mortem|acceptance-criteria|current as of → all 4 present
- grep drill-down|K-pass|≤3|rework → all 4 present
- grep cold-by-construction|did not author|kill channel → present
- grep v[0-9]+\.[0-9]+\.[0-9]+|SDD lines|through v → empty (W4 a,c)
- grep spec-reviewer|code-quality-reviewer → empty (W4 b)
- description cap: 1214 < 1536
- arch/SKILL.md: 124 lines < 600
- principle 09: 37/37 green
- render-rule-index --check: green
- principle 14: 1 pre-existing failure (broken refs in aif-docs/aif-skill-generator/aif-reference — NOT touched by this PR, verified by stash)

Prior-art: prior-art-evaluations.md#55 (Superpowers TDD-for-Skills paired-negative, verdict ADAPT — this smoke reuses that shipped structural-check pattern on a new surface (upstream skill references); no new engine, no dependency, no capability beyond the pattern already registered there).

§1.7: forward-check — both edited skills keep their doc-authority headers (.claude/skills/arch/SKILL.md:23, .claude/skills/night-mode/SKILL.md:6) per doc-authority-hierarchy.md §2-§3, and the new gate is deterministic vitest with zero API-billed calls per no-paid-llm-in-ci.md; backward-check sweep — the change class is «our SKILL.md files asserting upstream internals», and rather than checking only the two files this commit edits, packages/core/skills/upstream-skill-reference.test.ts:350 enumerates EVERY tracked .claude/skills/*/SKILL.md and resolves every superpowers: reference in it against the installed upstream (9/9 on the host, 3 real roots discovered), so sibling skills are swept by the mechanism itself rather than by assertion.

* fix(arch-v2-s-a): land W4(c) site-c paragraph fix (rework a7ab79417b95)

The previous commit c1fb91458 marked W4 done but left the night-mode/SKILL.md:29
paragraph still carrying the residual meta-instruction parenthetical
«(cite the behaviour, not a line range)” and the duplicate “SDD's BLOCKED handler”
that the rework review caught. This commit lands the one-line behaviour-only
citation the review verified: “SDD's BLOCKED handler, which re-dispatches with a
more capable model on the increment's own signal”.

Addresses blocking finding a7ab79417b95 (rework iteration 2/3): the fix was
present in the working tree but uncommitted, so `git diff origin/staging...HEAD`
did NOT carry it — a PR built from HEAD alone would still ship the [f784dd784a81]
defect.

Verification (in-repo, quoted per T3):
- grep -nE 'v[0-9]+\.[0-9]+\.[0-9]+|SDD lines|through v' .claude/skills/arch/SKILL.md .claude/skills/night-mode/SKILL.md → empty (W4 a/c clean)
- grep -n 'spec-reviewer\|code-quality-reviewer' .claude/skills/night-mode/SKILL.md → empty (W4 b clean)
- wc -l .claude/skills/night-mode/SKILL.md → 54 (< 600 gate)

Prior-art: skipped — 1-line fix to existing skill markdown, no new capability

* docs(arch-v2-s-a): preserve the acceptance report + draft the round-2 rework input

S-A round 1 (aif task 003b0678, branch feature/arch-v2-context-pipeline-s-a-003b06)
was judged REWORK by the Opus acceptance seat: 2 MAJOR + 7 MINOR, with
everything else accepted and the executor explicitly credited for
fabricating no upstream-side quote.

Two artefacts committed so the next session inherits the evidence rather
than the conclusion:

1. The acceptance report itself. It lived only in the session scratchpad
   under /private/tmp and would not have survived. It carries the three
   HOME-varied probes proving M2 (a present-but-broken upstream install
   reads as 'not installed'), the host-side upstream re-verification the
   container could not do, and the per-criterion table.

2. The round-2 rework kickoff, marked DRAFT and explicitly NOT
   dispatch-ready. Substantively complete — M1, M2, all seven MINORs,
   acceptance, descopes, park contract, traps — but two passes are owed
   first, both operator-flagged: (a) the cross-model edge is unaccounted
   for, since this input runs on GLM in aif rather than on Claude Code,
   and .claude/skills/claude-glm-executor-handoff/SKILL.md owns that
   contract; (b) check-kickoff-traps.sh and host-verify.sh have not been
   run against the file.

Round 1's own defect argues for (a): the M1 fix satisfied its grep on a
hyphen (code-quality reviewer vs code-quality-reviewer) — a literal
reading of an instruction, which is the divergence class a cross-model
prompt contract exists to pre-empt.

Prior-art: skipped — evidence record + draft dispatch input, no new capability

* fix(arch-v2-s-a): round-2 rework — 2 MAJOR + 7 MINOR from the acceptance report

Round 1 (aif task 003b0678) was judged REWORK by the Opus acceptance seat.
Round 2 was executed in-session rather than dispatched back to aif: the two
MAJORs were one sentence and one ~50-line function, against roughly a million
tokens for a full factory cycle, and the acceptance report had already
specified every fix — there was no plan left to make.

M1 — night-mode/SKILL.md:15. Round 1 repaired the opening clause but the same
sentence still assigned tiers to the roster it had just retired ("top-down
(spec/architecture) reviewer" / "bottom-up code-quality reviewer"). Upstream
SDD has one task reviewer covering both axes plus one broad final reviewer, so
there was no split to assign tiers to. Re-cast onto SDD's real seats. The
round-1 grep missed this because the survivor was spelled with one hyphen and
the check looked for two — criterion satisfied, drift alive. The round-2 grep
is deliberately wider (`code-quality[ -]reviewer`) and returns empty.

M2 — upstream-skill-reference.test.ts. The smoke could not tell "no upstream"
from "broken upstream": a version directory with no skills/ child produced a
verdict byte-identical to a machine with nothing installed, so a corrupted,
relocated, or permission-denied install passed the load-bearing check
silently. `discoverUpstreamRoots` now returns a structured result — roots,
baseFound, errors, searchedGlob — and the integration test fails loudly as
BROKEN when a base exists but yields no root, or when any readdir throws. The
marketplace segment is globbed rather than frozen to `superpowers-dev` (the
host carries two marketplaces), and the SKIPPED line names the globs searched
instead of HOME, which is the one field that cannot distinguish the two cases.
Three synthetic-HOME unit tests pin the discrimination.

MINORs: m1 comment described control flow the file does not have; m2
tautological assert replaced with a message-shape assert; m3 the cold
definition is now referenced at both consumer sites it claimed to gate; m4
upstream filename dropped (soft version pin); m5 folded into M2; m6 unique
filenames given as examples rather than asserted about absent prompts; m7
duplicate trigger/skip paragraph dropped and §1.5 renumbered.

Verified on the host: 9/9 upstream-skill-reference; probe (a) fails loudly
where round 1 passed silently; principles 09/12/14/15 green; rule-index green;
host-verify 4/4; check-kickoff-traps EXIT=0.

Prior-art: skipped — defect repair inside an existing test and two existing
skills; no new capability, dependency, or engine introduced.

* fix(arch-v2-s-a): absorb the cold code review — 1 BLOCKER + 6 MAJOR

The cold code-review seat returned REVISE on the round-2 diff. CI was green and
the fidelity seat had returned GO, which is exactly the gap T19 exists to cover:
form and WHAT-conformance were fine, the environment logic was not.

B1 (BLOCKER) — a healthy install behind an unreadable directory still read as
"not installed". `existsSync` returns false for BOTH "absent" and EACCES, so a
permission-denied install produced a silent SKIPPED pass — the precise
conflation this round set out to end, and one the file's own header forbade in
so many words. Discovery now probes with a `statSync` helper that keeps the
three answers distinct (absent / directory / unreadable) and records anything
that is not ENOENT. Reproduced before and after: chmod 000 on the marketplace
directory went from "9 passed, SKIPPED" to a loud BROKEN failure.

M1 + m2 — `Dirent.isDirectory()` is false for a symlinked directory, so a
symlinked marketplace or version directory hid a healthy install (and a
symlinked version directory produced a false BROKEN on a healthy one). The same
`statSync` probe follows symlinks, which is what the advertised
`cache/*/superpowers/*/skills` glob always implied.

M3 — `checkReferences` swallowed a readdir throw on a skills root, contradicting
this round's own contract that a throw is surfaced rather than swallowed. An
unreadable root is now its own named failure, so the diagnosis stops blaming the
reference for an IO problem (m3).

M4 — the M1 repair introduced a fresh contradiction with upstream: it assigned
the final whole-branch review to the CHEAPER tier, while SDD 6.2.0 SKILL.md:165
dispatches that review "on the most capable available model, not the session
default". Moved to the top tier. The acceptance report had suggested the wording;
per the round's own T-SA-R2-A the report is evidence, not authority.

M5 — the same drift (b) survived in four more places (`dual-reviewer`,
"two independent reviewers") that neither acceptance grep reached. Fixing only
the grepped site would have repeated round 1's exact failure.

M6 — the m3 cold-reference repair over-claimed: it said every channel from the
distillate onward is judged by a cold seat, but §1.5 step 3 has the verifier seat
distil and rule on its own distillate. Now states which channels are cold and why
the K-pass exists.

Deliberately NOT fixed, surfaced instead: orphaned cached versions counted as
"installed" (a new work item — the header now states the weaker claim and its
falsifier rather than pretending); coverage excluding shipped
skills/getff/references/*.md; two pre-existing arch/SKILL.md nits.

Verified on the host: 13/13 on the smoke, 66/66 with principles 09/12/14/15,
rule-index green, both round-2 greps still empty, arch 120 lines.

Prior-art: skipped — defect repair inside an existing test and two existing skills; no new capability, dependency, or engine introduced.

* fix(arch-v2-s-a): absorb the verification review — a regression my own fix introduced

The adversarial re-review confirmed the previous round's fixes hold (permission
denial is now loud at all seven levels; all four symlink shapes are followed;
10 of 13 mutations killed) and found that the fix itself had introduced a new
false alarm, plus one test that did not test what it named.

MAJOR — a stray FILE in the plugins cache turned a healthy install into a loud
BROKEN, and worse, turned "nothing installed" into a FAIL. Probing
`<entry>/superpowers` without first probing `<entry>` reports ENOTDIR for any
plain file, and a single `.DS_Store` — one Finder visit — was enough. The root
cause was treating every non-ENOENT errno as an alarm, so the fix is a
distinction rather than another special case: an error now means "we were
DENIED the answer", never "the answer is no". EACCES/EPERM can conceal an
install and stay loud; ENOENT/ENOTDIR/ELOOP are definitive negatives and are
skipped. That resolves the same class at every level at once, including the
ELOOP false alarm on a self-symlink.

MAJOR — the symlink test named "marketplace AND version dirs" but symlinked
only the marketplace, which resolves through ordinary path resolution; a
faithful lstatSync mutation at the version and skills probes left it green. Now
one case per level, each also resolving a reference through the discovered root
so the root is proven usable rather than merely listed. The same mutation now
kills four tests.

MINOR — an unreadable skill directory under a readable root (mode 0444) blamed
the reference for a permission problem; the root is now named. A bare `return`
guard reported a green tick for a test that asserted nothing under root —
`it.skipIf` reports a skip. An unset HOME was a silent SKIPPED pass; "we cannot
work out where to look" is not "nothing is installed".

MAJOR (docs) — arch/SKILL.md's description and its §2 bullets assigned tiers to
the two review seats, while the seat-instantiation paragraph puts BOTH seats on
the mid tier whenever the authoring session is top tier — which §0 says it
always is. The advertised tiers could therefore never occur. Seats are now
altitudes only; tiers are stated once, where they are decided.

MINOR (docs) — "brainstorming dispatches an author-side spec-document reviewer"
overstated upstream: the prompt template ships, but nothing in the 6.2.0 flow
references it (grep over the installed skills returns no referencing line) and
that flow's spec pass is a self-review. Replacing a negative-existence overclaim
with a positive one would have repeated W4(a) in mirror image.

Surfaced, not fixed (outside the round-2 permitted file set): the same W4(b)
roster drift survives in pipeline/SKILL.md:320-323, which additionally credits
it to SSOT #64 vocabulary.

Verified on the host: 18/18 on the smoke, 71/71 with principles 09/12/14/15;
`.DS_Store` with and without an install both behave; the permission case stays
loud; the lstatSync mutation kills 4 tests; rule-index green; arch 122 lines,
description 1216 chars.

Prior-art: skipped — defect repair inside an existing test and one existing skill; no new capability, dependency, or engine introduced.

* fix(arch-v2-s-a): drop a version pin my own W4(a) repair reintroduced

The narrow scope re-audit caught a third instance of the same shape that
produced M1: the sentence rewritten to stop overstating upstream ("dispatches"
→ "ships") reintroduced the literal version "6.2.0" at the exact site W4(a) had
cleaned. Round-1 W4 rejects version markers and round-1 §4 descopes upstream
version pins anywhere in the diff, but the acceptance grep requires a leading
`v`, so a bare semver passes it — criterion satisfied, drift alive, for the
third time in this stage.

Stated version-free; the claim needs no version to stand. Checked with the
widened grep the criterion should have used: `grep -nE '[0-9]+\.[0-9]+\.[0-9]+'`
over both skills now returns empty, as do both original acceptance greps.

Prior-art: skipped — one-clause wording fix in an existing skill; no capability, dependency, or engine involved.

---------

Co-authored-by: Test <test@example.com>
artyhoo added a commit that referenced this pull request Aug 10, 2026
…provenance repair (principle-11 F1 staging red) (#1375)

## Summary

Two concerns, one invited scope: (1) NEW project skill `.claude/skills/reviewer/SKILL.md` — the interactive review-session protocol, until now living only in the operator's personal `~/.claude/commands/reviewer.md` where repo machinery cannot see or update it (the #1374 severity-contract change had to be hand-patched into it the same day — the incident this closes; in-repo, a project skill takes precedence over the same-named personal command per the documented skill-over-command rule). (2) SSOT entry #249 — provenance repair for `.claude/rules/effort-worthiness.md`: principle 11 F1 is red on staging because the #1374 squash rebuilt the introducing commit from a PR body that omitted the `Prior-art:` line (the #1094#1097 class); the verbatim-path SSOT match fixes F1 for every subsequent PR.

## Changes

- NEW `.claude/skills/reviewer/SKILL.md`: three modes + verification-vs-synthesis economy split adapted 1:1 from the operator command; verdict grammar bound to `reviewer-discipline.md` §6 (Failure-scenario, ESCALATED, notes lane, zero-finding legitimacy); explicit subordination to the cold agents it does not replace and an explicit not-a-registry-role note (seat-lifecycle.md §1 three-roles cut respected — no seat-lifecycle edit).
- `.claude/skills/arch/SKILL.md:91`: stale «pending as of 2026-08-10» claim about the operator's global `/reviewer` resolved (hand-apply done same day; in-repo invocations now load the project skill).
- `docs/meta-factory/prior-art-evaluations.md`: entry #249 (REFERENCE) — registers the §8 item-5 per-family consult (Conventional Comments, Google eng-practices, Bezos Type-1/2, CBR, WIP limits, ADR/spec-kit/Kiro) already folded into the rule's §4, and records the detector-disagreement root cause (pr-body-prior-art's diff detector calls new rule-markdown non-capability while F1 counts it) with a widening trigger.
- Deliberately NOT changed: `setup.d/10-skills.sh` — consumer delivery of the reviewer skill is routed to advisor-pattern §8 item 9 (consumer-delivery stage with its own review), with an env+ tier recommendation (pairs with arch/pipeline at that profile).

## Prior-art consult

Prior-art: skipped — project-internal interactive-review skill adapting the operator's own global /reviewer command into the repo; subordinates to reviewer-discipline.md §6 (severity contract SSOT) and to the cold agents it does not replace; no new capability, no packages/ code

- [x] PR range is non-capability (one new skill markdown + one SSOT row + one line edit; zero `packages/` files, zero dependency changes); the trailer line above is carried in the PR body so it survives the squash into the introducing commit (the exact #1094#1097 / #1374 lesson this PR also repairs).
- [x] SSOT touch: entry #249 appended (append-only register; capability-commit-author write access per the ownership table).
- [x] Cold `agents/capability-reuse-auditor.md` pass run before handoff (source-before-shape Layer B): verdict THIN-ADAPT → GO, 12-candidate overlap set, trailer↔body consistent per clause; its one notes-lane finding (sibling §6 digest cross-pointer) applied in-branch.

## Test plan

- [x] §1.7-свод lands in squash-body (`gh pr merge --squash --body "$(gh pr view <N> --json body -q .body)"`)
- [x] `npm run --prefix packages/core test:principles` — 41 files green locally after #249 (principle 11 F1 was the one red: 14/14 after; principles 09/14/15 cover the new skill dynamically, 46/46)
- [x] `bash scripts/check-skill-drift.sh` — PASS (0 errors); doc-authority hook smoke on the new file — exit 0
- [x] Pre-push full substance sweep green at push time
- [x] Manual smoke: `.claude/skills/reviewer/SKILL.md` paired-negative sections present (`## Without this skill` / `## With this skill`); frontmatter description carries concrete RU+EN triggers per skill-description-quality.md §2

## Provenance

n/a — dialog-invited repo-skill addition + CI-red repair; no stage kickoff, no dispatch substrate.

## Review findings

Cold capability-reuse-audit (agents/capability-reuse-auditor.md, dialogue-blind on the just-authored file + intended trailer): verdict THIN-ADAPT, overall GO. Overlap set = 12 candidates (SSOT #64/#231/#236/#249 + aif-review family; skills/agents siblings; upstream Superpowers requesting-code-review + SDD). Key clears: the PR #858 class does not recur (the SDD executor/dual-reviewer loop is not re-described; spawning subagents without approval is forbidden by the skill's Hard bounds); subordination lines verified per file:line for all four owners named in the header. One notes-lane finding, fixed same round per the §6 contract: the repo now holds two §6 digests (this skill + agents/reviewer-discipline.md) — a cross-pointer with a drift rule was added to the skill's See also. Coverage honestly partial (headers-only for the non-overlapping skill/agent tail; effort-worthiness.md body unread by the auditor).

## Fidelity verdict

FIDELITY: skipped — dialog-invited non-stage PR (repo-skill addition + staging F1 repair); no kickoff/spec basis to audit against; cold reuse-audit GO recorded under Review findings.

## Parked questions

- Consumer tier for the reviewer skill: recommendation env+ (contour surface, pairs with arch/pipeline in setup.d/10-skills.sh:120-122); decision + consumer-generic rewording (the skill's Origin references the operator's home path) belong to advisor-pattern §8 item 9 — recorded in the memory card as an item-9 input.
- SSOT #249's «Trigger to revisit» carries the detector-widening trigger: a second squash-trailer F1 incident → widen pr-body-prior-art's detector to F1's artifact classes (rule/skill/agent markdown).

## §1.7 Self-discipline check (REQUIRED if PR touches discipline-bearing files)

### §1.7 Forward-check applied

New skill checked against every active layer: channel selection per rule-enforcement-channel-selection.md — on-demand skill load at the review-ask trigger, not always-on (description triggers per skill-description-quality.md §2, RU+EN); doc-authority header present with subordination lines (principle 09 dynamic skill enumerator green, packages/core/principles/09-doc-authority-hierarchy.test.ts); paired-negative sections present per packages/core/principles/15-skill-paired-negative.test.ts:50; provenance per principle 11 F1 — the `Prior-art:` trailer lives in this PR body (squash-safe) AND the artifact has SSOT keyword coverage, while the F1 red this PR repairs is closed by the verbatim path in docs/meta-factory/prior-art-evaluations.md entry #249; language-discipline — repo artifact in English; seat-lifecycle NOT extended — the three-roles cut at .claude/rules/seat-lifecycle.md:42 is respected by an explicit not-a-registry-role subordination line instead of a paths: edit; source-before-shape §1 — SSOT + .claude/skills/ + agents/ grepped before the body was written, and Layer B (agents/capability-reuse-auditor.md) run before handoff with verdict THIN-ADAPT/GO.

### §1.7 Backward-check applied

Class = artifacts carrying the interactive-reviewer protocol or a §6 severity-contract digest. Surfaces enumerated (grep over .claude/skills/**, agents/**, setup.d/, plus the out-of-repo command): ~/.claude/commands/reviewer.md — out-of-repo sibling, hand-patched to §6 grammar 2026-08-10, now shadowed in-repo by this skill (skill-over-command precedence), SWEPT-CLEAN; agents/reviewer-discipline.md:37 — sibling run-moment §6 digest, GAP-FOUND (no cross-link between the two digests) → FIXED this PR (See-also drift rule in .claude/skills/reviewer/SKILL.md); .claude/skills/arch/SKILL.md:91 — stale «pending» claim about the global command, GAP-FOUND → FIXED this PR; .claude/rules/reviewer-discipline.md:56 — the operating SSOT itself, untouched by design (both digests subordinate to it); packages/core/templates/shared/skill-context/aif-review/SKILL.md + aif-orchestrator-discipline — shipped consumer surfaces, deliberately DEFERRED to advisor-pattern §8 item 9 (never a silent copy; ownership-table read-only for sessions); setup.d/10-skills.sh:108-127 — the delivery manifest, deliberately untouched, tier decision recorded as an item-9 input. No other surface in the class (agents/fidelity-auditor.md + agents/review-sidecar.md carry the cold-protocol grammar shipped by #1374, different altitude, already current).
artyhoo added a commit that referenced this pull request Aug 18, 2026
…ws, routing bindings (#1462)

* feat(arch): frontier pacing delta in §1 — ADAPT of grill-me/grilling (SSOT #253)

Two deltas over the wrapped brainstorming loop: batch prerequisite-settled
questions per round (dependent ones stay serial), and enumerate-before-done —
the dialogue closes only when every design decision is answered or an explicit
operator fork. Superpowers 6.2.0 verified to lack the mechanism (grep over the
whole installed plugin: 0 hits; brainstorming pins one-question-per-message).

Prior-art: prior-art-evaluations.md#253 (grill-me/grilling, ADAPT — only the tree/frontier mechanic transfers; recommendation-per-question and probe-don't-ask already exist as H1 + T8/T20).

* feat(arch): grilling becomes the questioning engine — SSOT #253 lifted ADAPT→ADOPT

Operator-ratified design session (D1-D4): the compressed frontier-pacing
paraphrase measurably lost upstream's non-blocking probe rule (same-day cold
review vs the raw upstream text), so /arch §1 now consumes the grilling skill
AS IS via the mattpocock-skills companion plugin (MIT, versioned, precedent
#64 brainstorming) and keeps only a thin binding: brainstorming collision
resolution, probe routing (T20/§1.5), AskUserQuestion as the round carrier
(added to allowed-tools), and the spec's live decision register as the tree
surface. Register format lands in the spec-template obligation; SSOT #253
revisit triggers gain a named recording surface + a vendor-copy fallback arm.

Prior-art: prior-art-evaluations.md#253 (grill-me/grilling, ADOPT — companion plugin consumed AS IS; paraphrase channel measured lossy, hence the lift from ADAPT).

* feat(arch): §2 no-rerank rule — the two altitudes are never merged into one list

Adopted from mattpocock code-review's two-axis separation (one axis must not
mask the other) during the 2026-08-17 plugin sweep; the §2 seats already
report independently, this pins that their findings are presented side by
side and never reranked across altitudes.

Prior-art: prior-art-evaluations.md#253 (mattpocock-skills plugin sweep; doc-only edit, no new capability).

* docs(arch-prep): three-stack skill harmonization — collision map + raw ownership idea

Prep-doc for a future /arch design session: enumerates all three skill
populations (ours 16, superpowers 6.2.0 14, mattpocock-skills 1.2.3 35 — all
35 read in full), maps collisions per capability area (sharpest: Matt tdd vs
SP TDD contradict on refactor placement and seam scoping; diagnosing-bugs vs
systematic-debugging claim the same trigger space), lists the six available
resolution mechanisms with two unknowns (per-skill disable, routing
precedence) as probes, and drafts a one-owner-per-area map plus a live
decision register the design session starts from.

Prior-art: prior-art-evaluations.md#253 (grilling ADOPT — this prep extends the same three-stack comparison to the full plugin; doc-only, no capability).

* docs(arch-prep): §1.5 dependency edges — collision risk weighted by our hard references

Measured map of what our machinery hard-references upstream (grep over
skills/rules/agents/CLAUDE.md/templates): SDD is the most-referenced upstream
and crosses the shipped axis (tier-home.md); requesting-code-review is the
highest-risk collision zone because dispatcher/harvest contracts name it while
Matt's code-review claims the same trigger space; TDD/debugging collisions
carry routing risk only (zero hard edges from us). Also: D-H4 recorded as
answered (parallel commit 09569a3 landed mid-session), §7 gains the
re-probe-before-edit note.

Prior-art: prior-art-evaluations.md#253 (same three-stack comparison; doc-only edit).

* docs(arch-prep): DeepWiki pass on both satellites + our-side thinning audit

Per-repo DeepWiki interrogation folded in: (a) the measured routable surface
is exactly 11 mattpocock skills (user-invoked ones never enter the router —
his collision policy is the user/model-invoked split, confirmed live in this
session's skill listing); (b) Matt's refactor-out-of-loop is a June-2026
behavioral measurement («agents essentially never performed it»), not
doctrine — D-H2 needs our own corpus check; (c) TDD edge CORRECTED: SDD's
implementer-prompt.md:36 says bare «TDD», so the collision is
transitive-contract grade, not routing-only; (d) superpowers documents
Project > Personal > Plugin per-skill shadowing — new mechanism 7, P1
narrowed. New §4.5: our 16 skills audited — nothing deletable, orchestrator
is the one THIN candidate (D-H9); D-H10 TDD shadow, D-H11 domain-modeling
pairing added to the register.

Prior-art: prior-art-evaluations.md#253 (same three-stack comparison; doc-only edit).

* docs(arch-prep): slash-only planning skills evaluated as adoption candidates + 4 raw ideas

Operator correction folded in: «no collision» ≠ «no value» — the user-invoked
planning skills get per-skill adopt/adapt verdicts (wayfinder ADAPT strongest;
to-tickets ADAPT mechanizable; to-spec one section; implement REJECT; triage
two residues). New §4.6 carries four raw ideas for the design session: (1) the
decision map as the multi-session layer over /arch — D4's register lifted to
wayfinder shape, map-location sub-fork included; (2) kickoff Blocked-by edges
with a pipeline-computed frontier; (3) seams-first Testing-seams slot in the
spec template, unlocking the seams half of D-H2; (4) glossary SSOT as a
term-ownership generated index — the CONTEXT.md-free adaptation that makes the
grilling+domain-modeling pairing adoptable (D-H11 re-opened from defer).
Register grows D-H12-D-H14.

Prior-art: prior-art-evaluations.md#253 (same three-stack comparison; doc-only edit).

* docs(arch): three-stack skill harmonization — design spec + continuation handoff

Interview phase complete (frontier empty): P-1..P-6 operator premises,
15-area ownership map ratified (D-H1), decision register D-H0..D-H16 with
falsifiers, mechanism set (prune script, CONTEXT.md rule+test, claim
reorder, Blocked-by frontier, seams slot, aif plugin), probe register
P1/P2a-c/P5-pending/P6, routed-work inventory for §3 exit routing.

Awaiting §2 cold two-altitude review (this session's next step).

* docs(arch): harmonization spec v2 — round-1 cold-review dispositions landed

Both §2 seats returned REVISE (9 + 8 findings). All round-triggering
findings repaired in place: §1 restated as two declared lanes (TD-F1);
prune radius narrowed to 2 machine-globally-justified items per the
operator's F7 answer + --check pre-push drift detector (TD-F2/F7);
D-H17 completes the ownership map to all 11 model-invocable skills
(TD-F3); setup run re-bucketed attended (TD-F4); D-H5 claim mechanics
specified with real machinery + P4 restored (TD-F5, B-M1/M2); §5.6
non-target named (B-M3); /vitest transfer dissolved (B-M4); D-H7/D-H8
counter statuses corrected (B-M5); D-H13 adopts incumbent 'Depends on'
spelling (B-M6). New: P-7 premise + D-H18 consumer-axis contour routed
out via chip. Full dispositions: §9 v2 entry.

* docs(arch): harmonization spec v3 + SSOT #253 counter arm + REJECT rows #254-257

Round-2 delta review (both seats REVISE; all round-1 closures confirmed):
- --check channel corrected: owner:'maintainer' section in the pre-push.ts
  section registry (the file ships to consumers but maintainer sections
  never compose on a consumer layout, fail-closed) — .husky/pre-push is
  an exec dispatcher with no sections (convergent TD/B finding).
- D-H16 build item DISSOLVED: aif container mounts the host
  ~/.claude/plugins read-only (docker-compose.override.yml), so the
  plugin is already visible in-container and the prune/--check cover it
  by construction (measured round-2).
- 'counter armed' made true instead of re-worded: D-H7/D-H8 arm +
  observation No.0 appended to SSOT #253; REJECT rows #254-257 added
  (Matt implement, ADR dir, severity-less review model, total-sweep
  pruning). Spec SS8 item 5 DONE in-session.
Dispositions: spec SS9 v3 entry.

* docs(arch): harmonization contour GO — round-3 record + routed in-session edits

Round 3 (targeted delta): both cold seats GO. Spec header → REVIEWED-GO;
§9 round-3 entry (one TD MINOR accepted as recorded limit: container
premise rests on untracked local docker-compose.override.yml — covered
by D-H16 falsifier).

Routed §8 item 2 small edits, per spec:
- arch/SKILL.md §1: Testing seams slot added to the spec-template
  obligation (D-H14; seams-first adopted WITHOUT Matt's refactor placement)
- ai-doc/SKILL.md: skill-authoring ownership note (standard=ours,
  process=SP writing-skills, writing-for-agents=REFERENCE)
- rule-tests/SKILL.md: tautological-test anti-pattern REFERENCE note
  (D-H2 transfer (b))

* docs(arch): close harmonization contour handoff — full tail executed

Review GO (3 rounds), exit routing done (3 chips + in-session edits),
SSOT appends landed. Handoff retained as closure record; residue =
operator actions (spec SS8 item 1) + chip-routed umbrellas.

* docs(arch): consumer-axis satellite harmonization — design v1 + round-3 handoff

Round-2 /arch contour (D-H18): interview closed, D-C1..D-C8 ratified with
falsifiers; three-class collision model (factory CI / install-time census /
informed consent); detect+declare+prescribe mechanism recorded. Cold review
and exit routing DEFERRED behind the operator-mandated round-3 top-down
creative re-examination (P-C3) — handoff written for the fresh session.

* docs(arch): harmonization round 3 — registers amended, injected-context bindings land

Round 3 (D-C8, operator-mandated P-C3) executed per the handoff's membrane
phase order. Operator-axis spec v4: D-H15 SUPERSEDED — the prune apparatus
(script / wizard / --check pre-push section / gate P5) dissolved, replaced
by CLAUDE.md routing bindings (repo section + ~/.claude/CLAUDE.md
machine-global half, written in-session with live operator approval) +
meta-kickoff.template.md binding line (D-H10 fallback promoted to primary);
D-H8 gains a frontmatter-neutering ladder step. Round-2 spec v2: D-C1
re-cut to the thin form (static census prose + known-pair presence check;
inventory-join engine not built), D-C9 fourth-stack admission boundary
added (knowledge-work trio stays on SSOT #235). Round-3 handoff closed
with the continuation-state staleness correction; keen-shannon merged in
(3ae6981) so both specs live on one branch.

* docs(arch): P7 recorded — fresh-session bindings probe 2/2 vs P2 baseline

Both P2-class triggers flip with the CLAUDE.md bindings in context
(headless claude -p, fresh sessions reading the worktree CLAUDE.md from
disk). Method finding recorded: in-session subagent probes are invalid
for mid-session binding edits — subagents inherit the parent's
session-start CLAUDE.md snapshot (measured via a failed in-session probe
plus its diagnostic follow-up).

* docs(arch): round-3 review R1 — both seats REVISE, dispositions landed

Convergent BLOCKER fixed: the meta-kickoff.template.md binding line
REMOVED — .claude/skills/pipeline/ ships to consumers via GETFF_SKILLS_ENV
(setup.d/lib.sh:59) at the default env profile, so carrier #3 breached the
operator-axis membrane while buying no coverage; its removal restores all
8 install fingerprints to the baseline blob. §5.1's «no mechanical channel
at all» premise corrected (config layer only; frontmatter + a possible
Skill-matched PreToolUse hook priced — P8 records the hook UNVERIFIED:
guide claims no Skill matcher, live harness observation contradicts). P7
restated honestly (1 measured flip + 1 post-only confirmation). SSOT #253/
#257 got dated supersession notes (no prune ever executed). Five residual
prune assertions re-cut. Consumer spec: population corrected — TWO shipped
cc-plugin rows (superpowers + ast-grep, the latter disabled on the
operator's own machine); presence check re-keyed on installed_plugins.json
+ enabledPlugins; D-C5/D-C6 aligned; class-2 own-skills half recorded as
prose-only limit. ESCALATED to operator: ast-grep shipping fate (ESC-1) +
the detection-wire fork (TD-M2/P8). ~/.claude/CLAUDE.md section relocated
to file end (orphaned AIF bullet restored to its heading).

* docs(arch): round-3 review R2 — residuals closed, operator answers landed

Both R2 seats REVISE with a convergent root cause: R1 edited the surfaces
findings argued FROM, not every surface repeating the claim. Closed: §1
premise re-cut to config-layer wording; §1 scope guard now names the
pre-round-3 routed edits as verified degrade-safe REFERENCEs; sixth prune
assertion re-cut (D-H16); handoff header unmerged label; D-H15
exclusivity hedge; consumer D-C1/§7 re-keyed on installed_plugins.json +
enabledPlugins; both §8 inventories carry the escalations. Operator
answers recorded live: ESC-1 → retro-census BOTH shipped rows, keep
ast-grep on a clean census; P8 → VERIFY via the settings.json hand-off
(§8 item 7). Review round cap (2 REVISE) reached — residual state
surfaced in §9 instead of a third cold round.

* docs(kickoffs): round-3 exit routing — two build umbrellas authored

consumer-satellite-contract (thin form: retro-census of BOTH manifest
rows per the answered ESC-1, D-C2 principle test, AGENTS.md.template
section + parity line, install-registry-keyed presence check) and
skill-harmonization-mechanisms (CONTEXT.md pointer-rule test, four-part
claim machinery closing probe P4, Depends-on frontier). Both carry
host-verify contracts and the PR-pause note: they become dispatchable
only when the spec branch merges to staging.

* docs(kickoffs): declare the effort-worthiness L0 rigor label on both round-3 kickoffs

Principle 40 (`packages/core/principles/40-kickoff-rigor-label.test.ts`) requires every
post-cutoff kickoff to carry a `Rigor label … L0 …` line with a legal value. Both
round-3 kickoffs were authored without it and failed the gate at push time.

- skill-harmonization-mechanisms → `build-and-verify`: all three surviving stages are
  factory-internal and reversible, each with a live RED/GREEN seam proof.
- consumer-satellite-contract → `research-grade`: S3/S4 touch consumer-shipped
  surfaces (AGENTS.md.template, ./setup), which effort-worthiness §1 reserves for the
  research-grade contour.

Prior-art: skipped — mechanical gate compliance on two doc files, no new capability

* docs(arch): P8 CONFIRMED live — detection wire v0 declared; kickoff L0 labels

The operator-registered log-only PreToolUse Skill hook fired on a forced
model-invoked skill in a fresh headless session (JSON with tool_name=Skill
+ the skill name in tool_input). The guide-agent's 'skill loading bypasses
the tool pipeline' claim is falsified — the P6 failure class again.
Measured boundary: user-typed slash commands bypass the Skill tool
(invisible to the wire, irrelevant: misroutes are model-invocations).
TD-M2 closes — the log IS the v0 misroute detection wire feeding the D-H7
counter; spec §6 P8 + D-H8 + §8 item 7 updated. Both kickoffs gained
effort-worthiness L0 rigor labels (parallel session's edit kept as-is).

* test(install-sh): regenerate baselines for the three edited shipped skills

The branch edits `.claude/skills/{ai-doc,arch,rule-tests}/SKILL.md` — all three are
shipped artefacts, so their fingerprints move in every stack baseline that carries
them. Captured with `SNAPSHOT_MODE=capture bash tests/install-sh/snapshot.sh`.

Diff reviewed before committing (round-3 R1 precedent: an unreviewed template edit
turned 8 fingerprints stale): exactly three payload paths changed hash — rule-tests
(22 occurrences), arch (16), ai-doc (16) — and every fingerprint file is 1:1 on line
count, so no payload entered or left any stack.

Prior-art: skipped — snapshot regeneration after a shipped-file edit, no new capability

---------

Co-authored-by: Test <test@example.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant