chore(skills+agents): cleanup batch from 2026-05-16 skills+agents audit - #64
Merged
Merged
Conversation
The Step-2 graceful-degradation block in agents/docs-auditor.md echoed 'See references/self-testing-docs.md for the pattern.' — referencing a path that does not exist (agents/references/ absent in source + consumer). Symmetric to former line-12 instance fixed in c8fc153. Surfaced by 2026-05-16 Wave 2 audit (D-AuditC-1 R-1).
… D-AuditC-6 Two consumer-facing agents reference paths populated only by the AIF installer in consumer projects (scripts/audit-ai-docs.sh for docs-auditor; .ai-factory/RULES.md for best-practices-sidecar). In the source repo these paths are absent; the agents handle this via graceful degradation. Note the expectation explicitly in each agent's description, and add a CLAUDE.md Artifact Ownership Contract row clarifying ownership. Per decisions.md D-AuditC-6 (2026-05-16). Wave 2 R-6 + R-7 + ownership row.
Wave 5.3 hook implementation has shipped. Tense correction only. Per decisions.md D-AuditC-1 R-3 (2026-05-16).
Add a See-also erratum pointing to 2026-05-16-think-time-s17-gate-correction.md. H2 vs H10 architectural re-evaluation with corrected Stop semantics is deferred to implementation moment (Phase 11+); maintainer flagged residual doubt — flag deliberately preserved, not silently absorbed. Per decisions.md D6 (2026-05-16).
This was referenced May 16, 2026
artyhoo
added a commit
that referenced
this pull request
May 22, 2026
…n sequencing plan (#160) Records the autonomous Queue-mode N8 R-phase outcome into the sequencing plan §0 snapshot: - §0 N8 row: 🟡 «plan committed; impl pending» → ✅ R-phase DONE (#158→staging, Worker→Reviewer GO→anti-collusion), A-phase 🔲 pending D1/D2/D3 - §0 «what remains»: N8 A-phase gated on maintainer D1 (cost-lever A/B/C) / D2 (local-model bench) / D3 (offload priority), surfaced in findings §7; SSOT #64–#68 ship with capability - §2 Track-1 1.1 + §4 first-launch: marked DONE (recommendation actioned) No status invented — reconciled against PR #158 + the n8-rphase-findings deliverable.
This was referenced May 22, 2026
artyhoo
added a commit
that referenced
this pull request
May 22, 2026
#165) * docs(meta-factory): N8 A-phase gate — list D1/D2/D3 explicitly in Track-1 row + dep diagram (§7 link) Track-1 row 1.3 and the dependency diagram showed only Depends-on 1.1; the D1/D2/D3 maintainer gate lived only in the §0 prose snapshot. An orchestrator picking work from the operational table/diagram could start A-phase impl before the cost-lever / bench / priority decisions are made. Now both operational touchpoints name D1·D2·D3 + link findings §7. Prior-art: skipped — doc-only edit, no new capability (sequencing-plan prose, no code/dep/artifact added). * docs(meta-factory): clarify N5/N6b readiness + SSOT-ID collision in §0 (anti-confusion) Sessions kept reading 🔲 as 'blocked' regardless of cause. Disambiguate §0: - N5: 🔲 **BLOCKED** — gated on N7 live-dogfood trial (decision=C done, repo-side applied, live trial pending) + N2; can't give back before dogfooding reveals what's unique. - N6b: 🔲 **NOT blocked** — deps met (N3 portable TS-core = packages/core/hooks/pre-push.ts+checks/; the 7 .claude/hooks/*.sh are thin CC-event glue, not engine; + N6a ✅). Sequenced 'last' by choice, startable anytime. - 🔲 legend added: blocked (unmet dep) vs deprioritized (deps met, startable) — the exact distinction sessions confuse. - ⚠ SSOT-ID collision flagged: N7 took #64/#65; N8 findings §7 proposed #64–#68 → N8 A-phase renumbers to #66+ at admission. N7 row left untouched (freshly applied by its owning session). Status reconciled against origin/staging + Wave-10 TS file inventory.
Merged
6 tasks
artyhoo
added a commit
that referenced
this pull request
May 22, 2026
…E Superpowers + SSOT #64/#65 + retention=A coexist (#166) N7 (DECISION=C) repo-side application. Source-verified the Superpowers adoption against shipped SKILL.md (dual-channel DeepWiki + raw WebFetch, 2026-05-22) instead of assuming from name. - .claude/rules/parallel-subwave-isolation.md §4: drop the homegrown AST-detection build-target; REFERENCE Superpowers `using-git-worktrees` (Step 0 GIT_DIR!=GIT_COMMON skip → compatible with isolation:"worktree"). Stays Class C; new §5 §1.7 forward/backward note. - prior-art-evaluations.md: append SSOT #64 (SDD, ADOPT process-layer) + #65 (using-git-worktrees, ADOPT + REFERENCE). NOT #60/#61 — those were taken by the parallel channel-selection wave (the exact ID-collision §6 warned about). "#27 = Git-Isolation" was a misattribution (#27 is HANDOFF_MODE) — corrected, void #27-update skipped. - N7 patch §9 closure: source-verification findings, retention verdict A (coexist), ID corrections, classifier-blocked items. - roadmap banner reconciled (DECISION=C set, original framing preserved append-only); sequencing-plan N7 status updated. Retention=A: SDD review is post-implementation, orchestrator Phase -1 is pre-dispatch + owns quota/Mode/bootstrap meta-layer SDD lacks → coexist. Blocked (maintainer-applied): global Superpowers install + orchestrator- skill annotation — both denied by the auto-mode classifier (untrusted external code / self-modification). Live-dogfood trial pending those. §1.7: Forward — demotion complies with build-first-reuse-default.md (REFERENCE-over-BUILD) + no-paid-llm-in-ci.md (using-git-worktrees pure-git); Backward — scope-reducing, only edited bullet is parallel-subwave-isolation.md:37 (§4), SSOT cross-ref at prior-art-evaluations.md row #65, full self-reflexive note at parallel-subwave-isolation.md:46 Prior-art: skipped — N7 process-layer dogfood adoption; SSOT #64/#65 appended documenting ADOPT verdicts, no new dependency or capability code added
This was referenced May 25, 2026
artyhoo
added a commit
that referenced
this pull request
Jun 3, 2026
… (Stage 1) (#403) Adds .claude/skills/dispatcher/SKILL.md (~195 LOC) wiring 4 existing CLI primitives (dispatch.ts / harvest.ts / questions.ts / answer.ts) into a dispatch→monitor→Q&A→harvest→Phase-1→stage-gate→advance loop. DN-A = dual (CC-native primary + portable fallback via capability-check, not brand-name; @dual-pair: dispatcher-skill). DN-B = Option A behavior: technical forks resolved autonomously via superpowers:brainstorming on CC; strategic forks surfaced to operator via questions.ts. Discrimination discipline baked into prose (ask-question-reminder.sh is operator-internal, not shipped). Prior-art: prior-art-evaluations.md#111 (BUILD — no upstream covers cross-boundary dispatch→monitor→Q&A→harvest→stage-gate loop; SSOT consult + DeepWiki ≥3 + WebSearch ≥3 in rphase §c; T16 verified vs #64/#67/#88). §1.7: forward-check applied — BFR=BUILD via all 5 layers (SSOT consult + DeepWiki≥3 + WebSearch≥3, dispatcher-skill-rphase.md:223), dual not cc-only per dual-implementation-discipline.md:3, zero-paid-LLM per no-paid-llm-in-ci.md:1, T16 problem-class X-vs-Y at SKILL.md:189; backward-check sweep — complementary to pipeline/SKILL.md:313 (no supersession), consumes runtime-bridge primitives + park.ts #109, SSOT #111 appended at prior-art-evaluations.md:181.
artyhoo
added a commit
that referenced
this pull request
Jul 2, 2026
…-development (#859) BFR self-correction (operator-flagged #parallel-evolution-creep): the first night-mode re-described SDD's coordinator + implementer + dual-reviewer loop (~70% of SSOT #64). Slimmed from ~160 to ~44 lines — the loop is now DELEGATED to superpowers:subagent-driven-development, and the skill owns ONLY the overnight delta SDD lacks: unattended autonomy/fork policy, quota/backoff resilience (ADAPT AIF watchdogs #45), Workflow context-economy, verification discipline, the verified diff-visibility harness fact, and the unsupervised terminal condition. Header + paired-negative retained; principles 09/14/15 green. Prior-art: prior-art-evaluations.md#64 (Superpowers subagent-driven-development, ADOPT — night-mode now explicitly a thin overnight adapter over the SDD loop, not a re-implementation) + prior-art-evaluations.md#45 (AIF watchdogs self-healing, ADAPT — the quota-backoff resilience basis). Co-authored-by: t <t@t.co>
artyhoo
added a commit
that referenced
this pull request
Jul 3, 2026
…capability-reuse-auditor (#863) Ships **source-before-shape** — an edit-time discipline catching two recurring AI-laziness failures at authoring time, closing a recursive-self-application gap in the project's own operating rules. Origin (2026-07-02, operator-confirmed cross-session recurrence = promotion trigger): 1. BFR-reinvention — PR #858 shipped a night-mode SKILL.md re-describing the loop SSOT #64 owns; its trailer said "ADAPT #64" while the body re-described it, and it passed principle 11 F1 (which checks trailer presence, not reuse substance). #consult-as-trailer-not-input. 2. scope-from-memory — a launch prompt scoped from recall, not the spec (D1 → B). #claim-from-memory-not-source. Mechanism (judgment → injection, not a gate): - Layer A: .claude/rules/source-before-shape.md carries globs/inject markers → the existing inject-matching-rule.sh surfaces the reminder at edit-time (.claude/skills/**, agents/**, .claude/orchestrator-prompts/**). REUSE, zero new engine. Honestly disclosed once-per-session limitation (best-effort first-touch nudge). - Layer B: agents/capability-reuse-auditor.md — AI-agnostic overlap + trailer↔body auditor (no-paid-llm), doing the semantic pass F1 cannot. SSOT #196 (ADAPT, records the 6-item BFR-consult); principle 09 REQUIRED_HEADER_DOCS +2; install.sh SHIPPED_DOCS +1 + regenerated byte-identical baselines; AGENTS.md rule-index +1; self-reflection research-patch. Independently reviewed (3-lens Workflow: 2 approve + 1 adversarial revise applied). Verified: principles green, byte-identical 8/8, injection dogfooded live. Candidate T-trap #consult-as-trailer-not-input surfaced-not-applied to ai-laziness-traps.md (dedicated rule home instead). §1.7: Forward+backward in source-before-shape.md §6 — complies with no-paid-llm-in-ci.md §1, build-first-reuse-default.md §4, phase-research-coverage.md §1.11, doc-authority-hierarchy.md §2-§3; registered at packages/core/principles/09-doc-authority-hierarchy.ts:54; self-applies (SSOT #196 consult drove the shape, not memory). Prior-art: prior-art-evaluations.md#196 (source-before-shape mechanism — ADAPT: REUSE inject-matching-rule.sh channel + BUILD the AI-agnostic capability-reuse-auditor; dedup-first ADOPT-VOCABULARY; code-clone tools REJECT on T16 problem-class miss; Superpowers writing-skills REFERENCE).
7 tasks
artyhoo
added a commit
that referenced
this pull request
Jul 18, 2026
…1032) * feat(skill): claude-glm-executor-handoff + advisor-consult sub-form Ships a thin-adapter skill for the Claude→GLM cross-model dispatch edge inside aif-handoff, plus an advisor-consult sub-form on the existing REPORT BLOCKER field, plus a research-patch provenance. - new .claude/skills/claude-glm-executor-handoff/SKILL.md (130 lines): 5 verified GLM-5.2 facts (D1-D5, source-grounded to Z.ai docs) + 3 refuted-folklore entries + 6-block input contract + status translation + recovery protocol + §5 honest-gaps marker (5 designed- not-proven claims). Paired-negative block per principle 15. Subordinates to night-mode/SDD/orchestrator-worker-discipline. - new docs/meta-factory/research-patches/2026-07-18-claude-glm- executor-handoff-facts.md (89 lines): provenance for the skill. Verified-facts table + refuted-folklore log + probe design (P1/P2/P3). Substrate (P4) resolved same-day: aif-handoff supports per-agent model: frontmatter natively. - agents/orchestrator-worker-discipline.md (+50 lines): new Advisor- consult protocol section. Adds BLOCKER: advisor-consult: prefix sub-form on the existing REPORT schema (no schema extension). Coordinator routing table for 3 cases. Caps at 2 per task. Subordinates to night-mode delta 7 for overnight cases. - .claude/skills/night-mode/SKILL.md (+2 lines): one-line pointer to the new skill, anchored on the existing Harness portability paragraph. Prior-art: SSOT #64 (SDD, ADOPT — the loop), SSOT #201 (Anthropic Advisor tool, ADAPT — the strategy generalised from night-mode delta 7 for non-overnight dispatch), SSOT #200 (harness-config emission, referenced in skill D6). §1.7: forward+backward applied — agents/orchestrator-worker-discipline.md:51 introduces Advisor-consult protocol (new discipline sub-form); .claude/skills/claude-glm-executor-handoff/SKILL.md:99 carries the §5 honest-gaps marker gating contract-section promotion; backward sweep of .claude/skills/**/SKILL.md and agents/*.md confirmed no re-description of night-mode/SDD/orchestrator-worker-discipline surfaces (Layer B capability-reuse-auditor verdict THIN-ADAPT). Honest gaps: skill §5 lists 5 designed-not-proven claims; advisor-consult reliability is similarly under trial. Promotion to hardened contract requires live probes recorded as follow-up research patches. Two-round review: Reviewer A (concept, bottom-up) + Reviewer B (execution, top-down) on draft v2, then cold /reviewer Mode 3 pass on final state. All 4 BLOCKERs from Reviewer B applied (status enum DONE|BLOCKED|PARTIAL; pay-as-you-go endpoint /api/paas/v4/; cached pricing cut; 5 smuggled-unverified claims sourced). Layer B capability-reuse-auditor verdict: THIN-ADAPT (body genuinely subordinates, no re-description). * chore(synth-bundle): rebuild after skill addition Pre-push hook caught drift — claude-glm-executor-handoff/SKILL.md addition changed the synth-and-wire bundle output. Rebuild to match. * chore(backends): npm capability-matrix eslint freshness — 9.39.4 Pre-push freshness gate caught drift: matrix claimed eslint 10.4.0, locally-resolving eslint is 9.39.4. Live firing test passes on 9.39.4 (5/5) — diagnostic format stable across the version range. Unrelated to the GLM handoff PR; bundled to unblock the push. * revert(backends): npm capability-matrix back to eslint 10.4.0 (CI truth) Previous commit 8c3e636 set the matrix to 9.39.4 to match locally- installed eslint, but CI does a fresh install and resolves 10.4.0. Pre-push freshness gate passed locally but failed in CI. Root cause: local node_modules was stale; the matrix was correct as originally committed (10.4.0 matches what CI sees). Fix: bring local eslint in sync (npm install eslint@10.4.0) AND keep the matrix at 10.4.0. * chore(synth-bundle): rebuild after npm install eslint@10.4.0 --------- Co-authored-by: t <t@t.co>
artyhoo
added a commit
that referenced
this pull request
Jul 30, 2026
…tpocock/skills + superpowers (#1183) Answers the operator's «кажется мы опять переизобрели велосипед» for the orchestration contour (/arch, /pipeline, /dispatcher, night-mode, aif-doctor, packages/runtime-bridge), closing the explicit gap PR #1181's raw material left open (it read DeepWiki summaries, never the shipped SKILL.md bodies). Answer: no for the substrate — the runtime is lee-to/aif-handoff (AifHandoffBackend.ts:1-2 "adapter for the lee-to/aif-handoff runtime"; 0 of 2175 TS lines implement a runtime), the executor loop is superpowers SDD (SSOT #64), the ideation loop is superpowers brainstorming, and disable-model-invocation is convergent with mattpocock's primary mechanism. Yes, narrowly, for two claims in our own bodies: - arch/SKILL.md:79 asserts "no design-review skill exists" upstream, but skills/brainstorming/spec-document-reviewer-prompt.md ships in 6.1.1 AND 6.2.0 (orphaned, referenced by no skill body), and its loop was retired at v5.0.6 after an A/B finding "identical quality scores" for ~25 min overhead. - night-mode/SKILL.md:15,24 describe SDD's 5.1.0 two-reviewer roster (retired into one task-reviewer-prompt.md), overstate the whole-work-review gap (SDD:391-414 has a whole-branch Final Review), and silently cap rework at ~4 rounds against SDD:320's five. Both are recorded, not fixed — §5 scopes them as proposals (F2/F3); skills and rules are out of scope per the Artifact Ownership Contract. 10 per-capability verdicts on both BFR axes with a T16 problem-class line and a falsifier each; three retain BUILD, one flips to WATCHLIST because SSOT #111's search missed builderz-labs/mission-control (5.9k stars, created 2026-02-13 — before that evaluation). Method: SSOT consult first (18 rows cited by ID), DeepWiki x4 across 2 repos, WebSearch x3 phrasings, gh api + direct reads of 3 cached superpowers versions. context7 excluded per BFR §3 tooling caveat. Prior-art: prior-art-evaluations.md#230 (mattpocock/skills, verdict REFERENCE — per-capability KEEP NARROW/no-match, one already-convergent mechanism), added in this commit. Prior-art: prior-art-evaluations.md#231 (superpowers retired spec-document-review loop, verdict REFERENCE — standing negative evidence against /arch §2), added in this commit. Prior-art: prior-art-evaluations.md#232 (builderz-labs/mission-control, verdict WATCHLIST — fires SSOT #111's own revisit trigger), added in this commit. Co-authored-by: Test <test@example.com>
artyhoo
added a commit
that referenced
this pull request
Jul 31, 2026
…mbrella + S-A kickoffs (#1189) * docs(handoff): /arch v2 + context-pipeline session handoff — decisions, in-flight aif tasks, continuation protocol Prior-art: skipped — session handoff document only, no new capability; decisions it records cite their own SSOT rows (#231, #207, #179, #64). * docs(spec): /arch v2 + context-pipeline system design — layer model, pipeline arc, ADR-1..8 Step 5 of the 2026-07-31 handoff §4 protocol. Fable design authored on the Opus research distillate (spot-checked, freshness-barred) and the Opus cold critique (GO-WITH-PATCHES). All three critique blockers absorbed by re-derivation: #231 over-read retracted (ADR-5), 5/5-K1 incident count corrected to 2/5 and the primary/background split dropped (ADR-6), the calibration falsifier given an oracle via shadow-A/B + pre-declared threshold (ADR-5). M1-M7 absorbed as design constraints (population table, bounded drill-down + distillate K-pass, K6 candidate/adjudicate split, gate-channel re-route to pre-push/CI, operationalized bet falsifier, L1/L2 boundary re-drawn, option spaces spanned). Prior-art: skipped — design spec only, no new capability shipped * docs(arch-v2): umbrella kickoff S-A..S-F + S-A stage-scoped dispatch input Execution plan for the /arch v2 + context-pipeline track, derived from the 2026-07-31 design spec (ADR-1..8) by the Opus plan-writing seat. Umbrella kickoff: stage table S-A..S-F with per-stage scope, dependencies, tier classification (justified against CLAUDE.md's fixed criteria), acceptance and implemented ADRs; dispatch protocol (4-arm in-flight probe, Phase -1 cold review, bridge-profile marker rule with the fidelity-verdict precondition quoted verbatim and re-verified at dispatch); calibration-ledger bootstrap (ADR-5/6/8) with the ADR-8 token instrument named; cross-umbrella dependency on token-audit S1 (S-E only, two gates: merged AND content-read). The bottom seat + shadow-A/B station is marked active from S-B merge onward — S-A predates the contract implementation and is covered only by Phase -1 plus its own acceptance commands. Stated, not papered over. Plan-writer objections (§4, per «who must write the plan cannot rubber-stamp the design»): O-1 three wrapper drifts, not two, and one mis-described — upstream ships brainstorming/spec-document-reviewer-prompt.md in 5.1.0/6.1.1/ 6.2.0, night-mode:15's SDD roster does not match upstream, night-mode:29 cites stale upstream line numbers; O-2 the skill-exists-by-name smoke catches none of them and skips silently off-host; O-3 ADR-8's token metric had no named instrument (aif task tokenTotal/costUsd, verified live); O-4 the ledger principle test is vacuous before 5 rows; O-5 S-D's tier is a function of S-C's verdict; O-6 the spec's marker condition drops CLAUDE.md's «produced by /arch». S-A kickoff is stage-scoped and self-contained (handoff decision 11): W1-W6 with concrete file targets and a verification command each, host-verify contract, descopes, §1.7 obligation in enumeration format, T-enumeration plus three domain traps. Prior-art: skipped — kickoff/plan docs only, no new capability * docs(arch-v2): preserve track evidence artifacts — distillate, corrected idea, cold critique The design spec cites these three as its evidence chain (distillate → corrected idea → GO-WITH-PATCHES critique); they lived only in the session scratchpad under /private/tmp, which does not survive a reboot. Committed verbatim as session artifacts of the 2026-07-31 protocol run. Prior-art: skipped — evidence-record docs only, no new capability * docs(spec): point evidence-chain citations at the committed artifact files Prior-art: skipped — link fix in a design doc, no new capability * docs(spec): absorb token-audit S1 acceptance — fresh N2 numbers, ADR-3 falsifier fired S1 (task c781e8a9, accepted 2026-07-31) measured the repo-owned always-on set at 29-39% of the observed ~100k session-start total, firing ADR-3's pre-registered falsifier: the budget gate's asserted quantity is re-scoped to the repo-owned share (explicitly labelled), the harness remainder routes to settings-recommendations, and the InstructionsLoaded verification task doubles as the measurement-extension probe. N2 updated to the fresher script-reproducible per-environment numbers (140,216 B host vs 118,374 B container), replacing the older channel-level A7 pair. Prior-art: skipped — design-doc correction on fresh measurement, no new capability --------- Co-authored-by: Test <test@example.com>
artyhoo
pushed a commit
that referenced
this pull request
Jul 31, 2026
…n fix introduced The adversarial re-review confirmed the previous round's fixes hold (permission denial is now loud at all seven levels; all four symlink shapes are followed; 10 of 13 mutations killed) and found that the fix itself had introduced a new false alarm, plus one test that did not test what it named. MAJOR — a stray FILE in the plugins cache turned a healthy install into a loud BROKEN, and worse, turned "nothing installed" into a FAIL. Probing `<entry>/superpowers` without first probing `<entry>` reports ENOTDIR for any plain file, and a single `.DS_Store` — one Finder visit — was enough. The root cause was treating every non-ENOENT errno as an alarm, so the fix is a distinction rather than another special case: an error now means "we were DENIED the answer", never "the answer is no". EACCES/EPERM can conceal an install and stay loud; ENOENT/ENOTDIR/ELOOP are definitive negatives and are skipped. That resolves the same class at every level at once, including the ELOOP false alarm on a self-symlink. MAJOR — the symlink test named "marketplace AND version dirs" but symlinked only the marketplace, which resolves through ordinary path resolution; a faithful lstatSync mutation at the version and skills probes left it green. Now one case per level, each also resolving a reference through the discovered root so the root is proven usable rather than merely listed. The same mutation now kills four tests. MINOR — an unreadable skill directory under a readable root (mode 0444) blamed the reference for a permission problem; the root is now named. A bare `return` guard reported a green tick for a test that asserted nothing under root — `it.skipIf` reports a skip. An unset HOME was a silent SKIPPED pass; "we cannot work out where to look" is not "nothing is installed". MAJOR (docs) — arch/SKILL.md's description and its §2 bullets assigned tiers to the two review seats, while the seat-instantiation paragraph puts BOTH seats on the mid tier whenever the authoring session is top tier — which §0 says it always is. The advertised tiers could therefore never occur. Seats are now altitudes only; tiers are stated once, where they are decided. MINOR (docs) — "brainstorming dispatches an author-side spec-document reviewer" overstated upstream: the prompt template ships, but nothing in the 6.2.0 flow references it (grep over the installed skills returns no referencing line) and that flow's spec pass is a self-review. Replacing a negative-existence overclaim with a positive one would have repeated W4(a) in mirror image. Surfaced, not fixed (outside the round-2 permitted file set): the same W4(b) roster drift survives in pipeline/SKILL.md:320-323, which additionally credits it to SSOT #64 vocabulary. Verified on the host: 18/18 on the smoke, 71/71 with principles 09/12/14/15; `.DS_Store` with and without an install both behave; the permission case stays loud; the lstatSync mutation kills 4 tests; rule-index green; arch 122 lines, description 1216 chars. Prior-art: skipped — defect repair inside an existing test and one existing skill; no new capability, dependency, or engine introduced.
13 tasks
artyhoo
added a commit
that referenced
this pull request
Jul 31, 2026
…acceptance (#1192) * feat(arch-v2-s-a): /arch v2 rewrite — research contour, membrane/K-pass, kill channels, W4 drift fixes, upstream-ref smoke S-A stage of the arch-v2-context-pipeline umbrella. Implements W1-W6 per the design spec §2 + ADR-4 + §4 item 1. W1 — §1.5 research contour: trigger + explicit skip, research-spec template (required pre-mortem + acceptance-criteria fields), execution with freshness bar, distillation with GO/rework/kill verdict, seats as relative tiers. W2 — membrane + K-pass + bounded drill-down (ADR-4): default consumption rule with bounded recourse (not isolation), K1/K2 pass on each distillate before consumption, drill-down capped at ≤3 per artifact with recording requirement. W3 — cold definition promoted to a named single statement («did not author AND did not receive authoring context»), referenced from §1.5 K-pass and §3 exit, not re-stated per section. Kill channels enumerated with cost ordering. W4 — three wrapper drifts repaired at interface level per handoff decision 6 (no version pins, no line numbers, no upstream internals): (a) arch/SKILL.md:79 — negative-existence claim → upstream capability + our delta (b) night-mode/SKILL.md:15 — roster contents → capability delegated to SDD (c) night-mode/SKILL.md:29 — line-number citation → behaviour citation W5 — packages/core/skills/upstream-skill-reference.test.ts smoke: asserts every superpowers:<name> reference in .claude/skills/*/SKILL.md resolves to an installed upstream skill directory. Honest scope — does NOT cover W4 drifts (stated in test header). Environment-aware: discovers upstream by glob, SKIPPED when absent. Paired-negative RED+GREEN+SKIPPED all exercised. W6 — unique-filenames dispatch contract for parallel subagents added to §2 (handoff decision 13; near-clobber incident 2026-07-31). Verification: - grep §1.5|pre-mortem|acceptance-criteria|current as of → all 4 present - grep drill-down|K-pass|≤3|rework → all 4 present - grep cold-by-construction|did not author|kill channel → present - grep v[0-9]+\.[0-9]+\.[0-9]+|SDD lines|through v → empty (W4 a,c) - grep spec-reviewer|code-quality-reviewer → empty (W4 b) - description cap: 1214 < 1536 - arch/SKILL.md: 124 lines < 600 - principle 09: 37/37 green - render-rule-index --check: green - principle 14: 1 pre-existing failure (broken refs in aif-docs/aif-skill-generator/aif-reference — NOT touched by this PR, verified by stash) Prior-art: prior-art-evaluations.md#55 (Superpowers TDD-for-Skills paired-negative, verdict ADAPT — this smoke reuses that shipped structural-check pattern on a new surface (upstream skill references); no new engine, no dependency, no capability beyond the pattern already registered there). §1.7: forward-check — both edited skills keep their doc-authority headers (.claude/skills/arch/SKILL.md:23, .claude/skills/night-mode/SKILL.md:6) per doc-authority-hierarchy.md §2-§3, and the new gate is deterministic vitest with zero API-billed calls per no-paid-llm-in-ci.md; backward-check sweep — the change class is «our SKILL.md files asserting upstream internals», and rather than checking only the two files this commit edits, packages/core/skills/upstream-skill-reference.test.ts:350 enumerates EVERY tracked .claude/skills/*/SKILL.md and resolves every superpowers: reference in it against the installed upstream (9/9 on the host, 3 real roots discovered), so sibling skills are swept by the mechanism itself rather than by assertion. * fix(arch-v2-s-a): land W4(c) site-c paragraph fix (rework a7ab79417b95) The previous commit c1fb91458 marked W4 done but left the night-mode/SKILL.md:29 paragraph still carrying the residual meta-instruction parenthetical «(cite the behaviour, not a line range)” and the duplicate “SDD's BLOCKED handler” that the rework review caught. This commit lands the one-line behaviour-only citation the review verified: “SDD's BLOCKED handler, which re-dispatches with a more capable model on the increment's own signal”. Addresses blocking finding a7ab79417b95 (rework iteration 2/3): the fix was present in the working tree but uncommitted, so `git diff origin/staging...HEAD` did NOT carry it — a PR built from HEAD alone would still ship the [f784dd784a81] defect. Verification (in-repo, quoted per T3): - grep -nE 'v[0-9]+\.[0-9]+\.[0-9]+|SDD lines|through v' .claude/skills/arch/SKILL.md .claude/skills/night-mode/SKILL.md → empty (W4 a/c clean) - grep -n 'spec-reviewer\|code-quality-reviewer' .claude/skills/night-mode/SKILL.md → empty (W4 b clean) - wc -l .claude/skills/night-mode/SKILL.md → 54 (< 600 gate) Prior-art: skipped — 1-line fix to existing skill markdown, no new capability * docs(arch-v2-s-a): preserve the acceptance report + draft the round-2 rework input S-A round 1 (aif task 003b0678, branch feature/arch-v2-context-pipeline-s-a-003b06) was judged REWORK by the Opus acceptance seat: 2 MAJOR + 7 MINOR, with everything else accepted and the executor explicitly credited for fabricating no upstream-side quote. Two artefacts committed so the next session inherits the evidence rather than the conclusion: 1. The acceptance report itself. It lived only in the session scratchpad under /private/tmp and would not have survived. It carries the three HOME-varied probes proving M2 (a present-but-broken upstream install reads as 'not installed'), the host-side upstream re-verification the container could not do, and the per-criterion table. 2. The round-2 rework kickoff, marked DRAFT and explicitly NOT dispatch-ready. Substantively complete — M1, M2, all seven MINORs, acceptance, descopes, park contract, traps — but two passes are owed first, both operator-flagged: (a) the cross-model edge is unaccounted for, since this input runs on GLM in aif rather than on Claude Code, and .claude/skills/claude-glm-executor-handoff/SKILL.md owns that contract; (b) check-kickoff-traps.sh and host-verify.sh have not been run against the file. Round 1's own defect argues for (a): the M1 fix satisfied its grep on a hyphen (code-quality reviewer vs code-quality-reviewer) — a literal reading of an instruction, which is the divergence class a cross-model prompt contract exists to pre-empt. Prior-art: skipped — evidence record + draft dispatch input, no new capability * fix(arch-v2-s-a): round-2 rework — 2 MAJOR + 7 MINOR from the acceptance report Round 1 (aif task 003b0678) was judged REWORK by the Opus acceptance seat. Round 2 was executed in-session rather than dispatched back to aif: the two MAJORs were one sentence and one ~50-line function, against roughly a million tokens for a full factory cycle, and the acceptance report had already specified every fix — there was no plan left to make. M1 — night-mode/SKILL.md:15. Round 1 repaired the opening clause but the same sentence still assigned tiers to the roster it had just retired ("top-down (spec/architecture) reviewer" / "bottom-up code-quality reviewer"). Upstream SDD has one task reviewer covering both axes plus one broad final reviewer, so there was no split to assign tiers to. Re-cast onto SDD's real seats. The round-1 grep missed this because the survivor was spelled with one hyphen and the check looked for two — criterion satisfied, drift alive. The round-2 grep is deliberately wider (`code-quality[ -]reviewer`) and returns empty. M2 — upstream-skill-reference.test.ts. The smoke could not tell "no upstream" from "broken upstream": a version directory with no skills/ child produced a verdict byte-identical to a machine with nothing installed, so a corrupted, relocated, or permission-denied install passed the load-bearing check silently. `discoverUpstreamRoots` now returns a structured result — roots, baseFound, errors, searchedGlob — and the integration test fails loudly as BROKEN when a base exists but yields no root, or when any readdir throws. The marketplace segment is globbed rather than frozen to `superpowers-dev` (the host carries two marketplaces), and the SKIPPED line names the globs searched instead of HOME, which is the one field that cannot distinguish the two cases. Three synthetic-HOME unit tests pin the discrimination. MINORs: m1 comment described control flow the file does not have; m2 tautological assert replaced with a message-shape assert; m3 the cold definition is now referenced at both consumer sites it claimed to gate; m4 upstream filename dropped (soft version pin); m5 folded into M2; m6 unique filenames given as examples rather than asserted about absent prompts; m7 duplicate trigger/skip paragraph dropped and §1.5 renumbered. Verified on the host: 9/9 upstream-skill-reference; probe (a) fails loudly where round 1 passed silently; principles 09/12/14/15 green; rule-index green; host-verify 4/4; check-kickoff-traps EXIT=0. Prior-art: skipped — defect repair inside an existing test and two existing skills; no new capability, dependency, or engine introduced. * fix(arch-v2-s-a): absorb the cold code review — 1 BLOCKER + 6 MAJOR The cold code-review seat returned REVISE on the round-2 diff. CI was green and the fidelity seat had returned GO, which is exactly the gap T19 exists to cover: form and WHAT-conformance were fine, the environment logic was not. B1 (BLOCKER) — a healthy install behind an unreadable directory still read as "not installed". `existsSync` returns false for BOTH "absent" and EACCES, so a permission-denied install produced a silent SKIPPED pass — the precise conflation this round set out to end, and one the file's own header forbade in so many words. Discovery now probes with a `statSync` helper that keeps the three answers distinct (absent / directory / unreadable) and records anything that is not ENOENT. Reproduced before and after: chmod 000 on the marketplace directory went from "9 passed, SKIPPED" to a loud BROKEN failure. M1 + m2 — `Dirent.isDirectory()` is false for a symlinked directory, so a symlinked marketplace or version directory hid a healthy install (and a symlinked version directory produced a false BROKEN on a healthy one). The same `statSync` probe follows symlinks, which is what the advertised `cache/*/superpowers/*/skills` glob always implied. M3 — `checkReferences` swallowed a readdir throw on a skills root, contradicting this round's own contract that a throw is surfaced rather than swallowed. An unreadable root is now its own named failure, so the diagnosis stops blaming the reference for an IO problem (m3). M4 — the M1 repair introduced a fresh contradiction with upstream: it assigned the final whole-branch review to the CHEAPER tier, while SDD 6.2.0 SKILL.md:165 dispatches that review "on the most capable available model, not the session default". Moved to the top tier. The acceptance report had suggested the wording; per the round's own T-SA-R2-A the report is evidence, not authority. M5 — the same drift (b) survived in four more places (`dual-reviewer`, "two independent reviewers") that neither acceptance grep reached. Fixing only the grepped site would have repeated round 1's exact failure. M6 — the m3 cold-reference repair over-claimed: it said every channel from the distillate onward is judged by a cold seat, but §1.5 step 3 has the verifier seat distil and rule on its own distillate. Now states which channels are cold and why the K-pass exists. Deliberately NOT fixed, surfaced instead: orphaned cached versions counted as "installed" (a new work item — the header now states the weaker claim and its falsifier rather than pretending); coverage excluding shipped skills/getff/references/*.md; two pre-existing arch/SKILL.md nits. Verified on the host: 13/13 on the smoke, 66/66 with principles 09/12/14/15, rule-index green, both round-2 greps still empty, arch 120 lines. Prior-art: skipped — defect repair inside an existing test and two existing skills; no new capability, dependency, or engine introduced. * fix(arch-v2-s-a): absorb the verification review — a regression my own fix introduced The adversarial re-review confirmed the previous round's fixes hold (permission denial is now loud at all seven levels; all four symlink shapes are followed; 10 of 13 mutations killed) and found that the fix itself had introduced a new false alarm, plus one test that did not test what it named. MAJOR — a stray FILE in the plugins cache turned a healthy install into a loud BROKEN, and worse, turned "nothing installed" into a FAIL. Probing `<entry>/superpowers` without first probing `<entry>` reports ENOTDIR for any plain file, and a single `.DS_Store` — one Finder visit — was enough. The root cause was treating every non-ENOENT errno as an alarm, so the fix is a distinction rather than another special case: an error now means "we were DENIED the answer", never "the answer is no". EACCES/EPERM can conceal an install and stay loud; ENOENT/ENOTDIR/ELOOP are definitive negatives and are skipped. That resolves the same class at every level at once, including the ELOOP false alarm on a self-symlink. MAJOR — the symlink test named "marketplace AND version dirs" but symlinked only the marketplace, which resolves through ordinary path resolution; a faithful lstatSync mutation at the version and skills probes left it green. Now one case per level, each also resolving a reference through the discovered root so the root is proven usable rather than merely listed. The same mutation now kills four tests. MINOR — an unreadable skill directory under a readable root (mode 0444) blamed the reference for a permission problem; the root is now named. A bare `return` guard reported a green tick for a test that asserted nothing under root — `it.skipIf` reports a skip. An unset HOME was a silent SKIPPED pass; "we cannot work out where to look" is not "nothing is installed". MAJOR (docs) — arch/SKILL.md's description and its §2 bullets assigned tiers to the two review seats, while the seat-instantiation paragraph puts BOTH seats on the mid tier whenever the authoring session is top tier — which §0 says it always is. The advertised tiers could therefore never occur. Seats are now altitudes only; tiers are stated once, where they are decided. MINOR (docs) — "brainstorming dispatches an author-side spec-document reviewer" overstated upstream: the prompt template ships, but nothing in the 6.2.0 flow references it (grep over the installed skills returns no referencing line) and that flow's spec pass is a self-review. Replacing a negative-existence overclaim with a positive one would have repeated W4(a) in mirror image. Surfaced, not fixed (outside the round-2 permitted file set): the same W4(b) roster drift survives in pipeline/SKILL.md:320-323, which additionally credits it to SSOT #64 vocabulary. Verified on the host: 18/18 on the smoke, 71/71 with principles 09/12/14/15; `.DS_Store` with and without an install both behave; the permission case stays loud; the lstatSync mutation kills 4 tests; rule-index green; arch 122 lines, description 1216 chars. Prior-art: skipped — defect repair inside an existing test and one existing skill; no new capability, dependency, or engine introduced. * fix(arch-v2-s-a): drop a version pin my own W4(a) repair reintroduced The narrow scope re-audit caught a third instance of the same shape that produced M1: the sentence rewritten to stop overstating upstream ("dispatches" → "ships") reintroduced the literal version "6.2.0" at the exact site W4(a) had cleaned. Round-1 W4 rejects version markers and round-1 §4 descopes upstream version pins anywhere in the diff, but the acceptance grep requires a leading `v`, so a bare semver passes it — criterion satisfied, drift alive, for the third time in this stage. Stated version-free; the claim needs no version to stand. Checked with the widened grep the criterion should have used: `grep -nE '[0-9]+\.[0-9]+\.[0-9]+'` over both skills now returns empty, as do both original acceptance greps. Prior-art: skipped — one-clause wording fix in an existing skill; no capability, dependency, or engine involved. --------- Co-authored-by: Test <test@example.com>
This was referenced Aug 1, 2026
Merged
Merged
8 tasks
artyhoo
added a commit
that referenced
this pull request
Aug 10, 2026
…provenance repair (principle-11 F1 staging red) (#1375) ## Summary Two concerns, one invited scope: (1) NEW project skill `.claude/skills/reviewer/SKILL.md` — the interactive review-session protocol, until now living only in the operator's personal `~/.claude/commands/reviewer.md` where repo machinery cannot see or update it (the #1374 severity-contract change had to be hand-patched into it the same day — the incident this closes; in-repo, a project skill takes precedence over the same-named personal command per the documented skill-over-command rule). (2) SSOT entry #249 — provenance repair for `.claude/rules/effort-worthiness.md`: principle 11 F1 is red on staging because the #1374 squash rebuilt the introducing commit from a PR body that omitted the `Prior-art:` line (the #1094→#1097 class); the verbatim-path SSOT match fixes F1 for every subsequent PR. ## Changes - NEW `.claude/skills/reviewer/SKILL.md`: three modes + verification-vs-synthesis economy split adapted 1:1 from the operator command; verdict grammar bound to `reviewer-discipline.md` §6 (Failure-scenario, ESCALATED, notes lane, zero-finding legitimacy); explicit subordination to the cold agents it does not replace and an explicit not-a-registry-role note (seat-lifecycle.md §1 three-roles cut respected — no seat-lifecycle edit). - `.claude/skills/arch/SKILL.md:91`: stale «pending as of 2026-08-10» claim about the operator's global `/reviewer` resolved (hand-apply done same day; in-repo invocations now load the project skill). - `docs/meta-factory/prior-art-evaluations.md`: entry #249 (REFERENCE) — registers the §8 item-5 per-family consult (Conventional Comments, Google eng-practices, Bezos Type-1/2, CBR, WIP limits, ADR/spec-kit/Kiro) already folded into the rule's §4, and records the detector-disagreement root cause (pr-body-prior-art's diff detector calls new rule-markdown non-capability while F1 counts it) with a widening trigger. - Deliberately NOT changed: `setup.d/10-skills.sh` — consumer delivery of the reviewer skill is routed to advisor-pattern §8 item 9 (consumer-delivery stage with its own review), with an env+ tier recommendation (pairs with arch/pipeline at that profile). ## Prior-art consult Prior-art: skipped — project-internal interactive-review skill adapting the operator's own global /reviewer command into the repo; subordinates to reviewer-discipline.md §6 (severity contract SSOT) and to the cold agents it does not replace; no new capability, no packages/ code - [x] PR range is non-capability (one new skill markdown + one SSOT row + one line edit; zero `packages/` files, zero dependency changes); the trailer line above is carried in the PR body so it survives the squash into the introducing commit (the exact #1094→#1097 / #1374 lesson this PR also repairs). - [x] SSOT touch: entry #249 appended (append-only register; capability-commit-author write access per the ownership table). - [x] Cold `agents/capability-reuse-auditor.md` pass run before handoff (source-before-shape Layer B): verdict THIN-ADAPT → GO, 12-candidate overlap set, trailer↔body consistent per clause; its one notes-lane finding (sibling §6 digest cross-pointer) applied in-branch. ## Test plan - [x] §1.7-свод lands in squash-body (`gh pr merge --squash --body "$(gh pr view <N> --json body -q .body)"`) - [x] `npm run --prefix packages/core test:principles` — 41 files green locally after #249 (principle 11 F1 was the one red: 14/14 after; principles 09/14/15 cover the new skill dynamically, 46/46) - [x] `bash scripts/check-skill-drift.sh` — PASS (0 errors); doc-authority hook smoke on the new file — exit 0 - [x] Pre-push full substance sweep green at push time - [x] Manual smoke: `.claude/skills/reviewer/SKILL.md` paired-negative sections present (`## Without this skill` / `## With this skill`); frontmatter description carries concrete RU+EN triggers per skill-description-quality.md §2 ## Provenance n/a — dialog-invited repo-skill addition + CI-red repair; no stage kickoff, no dispatch substrate. ## Review findings Cold capability-reuse-audit (agents/capability-reuse-auditor.md, dialogue-blind on the just-authored file + intended trailer): verdict THIN-ADAPT, overall GO. Overlap set = 12 candidates (SSOT #64/#231/#236/#249 + aif-review family; skills/agents siblings; upstream Superpowers requesting-code-review + SDD). Key clears: the PR #858 class does not recur (the SDD executor/dual-reviewer loop is not re-described; spawning subagents without approval is forbidden by the skill's Hard bounds); subordination lines verified per file:line for all four owners named in the header. One notes-lane finding, fixed same round per the §6 contract: the repo now holds two §6 digests (this skill + agents/reviewer-discipline.md) — a cross-pointer with a drift rule was added to the skill's See also. Coverage honestly partial (headers-only for the non-overlapping skill/agent tail; effort-worthiness.md body unread by the auditor). ## Fidelity verdict FIDELITY: skipped — dialog-invited non-stage PR (repo-skill addition + staging F1 repair); no kickoff/spec basis to audit against; cold reuse-audit GO recorded under Review findings. ## Parked questions - Consumer tier for the reviewer skill: recommendation env+ (contour surface, pairs with arch/pipeline in setup.d/10-skills.sh:120-122); decision + consumer-generic rewording (the skill's Origin references the operator's home path) belong to advisor-pattern §8 item 9 — recorded in the memory card as an item-9 input. - SSOT #249's «Trigger to revisit» carries the detector-widening trigger: a second squash-trailer F1 incident → widen pr-body-prior-art's detector to F1's artifact classes (rule/skill/agent markdown). ## §1.7 Self-discipline check (REQUIRED if PR touches discipline-bearing files) ### §1.7 Forward-check applied New skill checked against every active layer: channel selection per rule-enforcement-channel-selection.md — on-demand skill load at the review-ask trigger, not always-on (description triggers per skill-description-quality.md §2, RU+EN); doc-authority header present with subordination lines (principle 09 dynamic skill enumerator green, packages/core/principles/09-doc-authority-hierarchy.test.ts); paired-negative sections present per packages/core/principles/15-skill-paired-negative.test.ts:50; provenance per principle 11 F1 — the `Prior-art:` trailer lives in this PR body (squash-safe) AND the artifact has SSOT keyword coverage, while the F1 red this PR repairs is closed by the verbatim path in docs/meta-factory/prior-art-evaluations.md entry #249; language-discipline — repo artifact in English; seat-lifecycle NOT extended — the three-roles cut at .claude/rules/seat-lifecycle.md:42 is respected by an explicit not-a-registry-role subordination line instead of a paths: edit; source-before-shape §1 — SSOT + .claude/skills/ + agents/ grepped before the body was written, and Layer B (agents/capability-reuse-auditor.md) run before handoff with verdict THIN-ADAPT/GO. ### §1.7 Backward-check applied Class = artifacts carrying the interactive-reviewer protocol or a §6 severity-contract digest. Surfaces enumerated (grep over .claude/skills/**, agents/**, setup.d/, plus the out-of-repo command): ~/.claude/commands/reviewer.md — out-of-repo sibling, hand-patched to §6 grammar 2026-08-10, now shadowed in-repo by this skill (skill-over-command precedence), SWEPT-CLEAN; agents/reviewer-discipline.md:37 — sibling run-moment §6 digest, GAP-FOUND (no cross-link between the two digests) → FIXED this PR (See-also drift rule in .claude/skills/reviewer/SKILL.md); .claude/skills/arch/SKILL.md:91 — stale «pending» claim about the global command, GAP-FOUND → FIXED this PR; .claude/rules/reviewer-discipline.md:56 — the operating SSOT itself, untouched by design (both digests subordinate to it); packages/core/templates/shared/skill-context/aif-review/SKILL.md + aif-orchestrator-discipline — shipped consumer surfaces, deliberately DEFERRED to advisor-pattern §8 item 9 (never a silent copy; ownership-table read-only for sessions); setup.d/10-skills.sh:108-127 — the delivery manifest, deliberately untouched, tier decision recorded as an item-9 input. No other surface in the class (agents/fidelity-auditor.md + agents/review-sidecar.md carry the cold-protocol grammar shipped by #1374, different altitude, already current).
artyhoo
added a commit
that referenced
this pull request
Aug 18, 2026
…ws, routing bindings (#1462) * feat(arch): frontier pacing delta in §1 — ADAPT of grill-me/grilling (SSOT #253) Two deltas over the wrapped brainstorming loop: batch prerequisite-settled questions per round (dependent ones stay serial), and enumerate-before-done — the dialogue closes only when every design decision is answered or an explicit operator fork. Superpowers 6.2.0 verified to lack the mechanism (grep over the whole installed plugin: 0 hits; brainstorming pins one-question-per-message). Prior-art: prior-art-evaluations.md#253 (grill-me/grilling, ADAPT — only the tree/frontier mechanic transfers; recommendation-per-question and probe-don't-ask already exist as H1 + T8/T20). * feat(arch): grilling becomes the questioning engine — SSOT #253 lifted ADAPT→ADOPT Operator-ratified design session (D1-D4): the compressed frontier-pacing paraphrase measurably lost upstream's non-blocking probe rule (same-day cold review vs the raw upstream text), so /arch §1 now consumes the grilling skill AS IS via the mattpocock-skills companion plugin (MIT, versioned, precedent #64 brainstorming) and keeps only a thin binding: brainstorming collision resolution, probe routing (T20/§1.5), AskUserQuestion as the round carrier (added to allowed-tools), and the spec's live decision register as the tree surface. Register format lands in the spec-template obligation; SSOT #253 revisit triggers gain a named recording surface + a vendor-copy fallback arm. Prior-art: prior-art-evaluations.md#253 (grill-me/grilling, ADOPT — companion plugin consumed AS IS; paraphrase channel measured lossy, hence the lift from ADAPT). * feat(arch): §2 no-rerank rule — the two altitudes are never merged into one list Adopted from mattpocock code-review's two-axis separation (one axis must not mask the other) during the 2026-08-17 plugin sweep; the §2 seats already report independently, this pins that their findings are presented side by side and never reranked across altitudes. Prior-art: prior-art-evaluations.md#253 (mattpocock-skills plugin sweep; doc-only edit, no new capability). * docs(arch-prep): three-stack skill harmonization — collision map + raw ownership idea Prep-doc for a future /arch design session: enumerates all three skill populations (ours 16, superpowers 6.2.0 14, mattpocock-skills 1.2.3 35 — all 35 read in full), maps collisions per capability area (sharpest: Matt tdd vs SP TDD contradict on refactor placement and seam scoping; diagnosing-bugs vs systematic-debugging claim the same trigger space), lists the six available resolution mechanisms with two unknowns (per-skill disable, routing precedence) as probes, and drafts a one-owner-per-area map plus a live decision register the design session starts from. Prior-art: prior-art-evaluations.md#253 (grilling ADOPT — this prep extends the same three-stack comparison to the full plugin; doc-only, no capability). * docs(arch-prep): §1.5 dependency edges — collision risk weighted by our hard references Measured map of what our machinery hard-references upstream (grep over skills/rules/agents/CLAUDE.md/templates): SDD is the most-referenced upstream and crosses the shipped axis (tier-home.md); requesting-code-review is the highest-risk collision zone because dispatcher/harvest contracts name it while Matt's code-review claims the same trigger space; TDD/debugging collisions carry routing risk only (zero hard edges from us). Also: D-H4 recorded as answered (parallel commit 09569a3 landed mid-session), §7 gains the re-probe-before-edit note. Prior-art: prior-art-evaluations.md#253 (same three-stack comparison; doc-only edit). * docs(arch-prep): DeepWiki pass on both satellites + our-side thinning audit Per-repo DeepWiki interrogation folded in: (a) the measured routable surface is exactly 11 mattpocock skills (user-invoked ones never enter the router — his collision policy is the user/model-invoked split, confirmed live in this session's skill listing); (b) Matt's refactor-out-of-loop is a June-2026 behavioral measurement («agents essentially never performed it»), not doctrine — D-H2 needs our own corpus check; (c) TDD edge CORRECTED: SDD's implementer-prompt.md:36 says bare «TDD», so the collision is transitive-contract grade, not routing-only; (d) superpowers documents Project > Personal > Plugin per-skill shadowing — new mechanism 7, P1 narrowed. New §4.5: our 16 skills audited — nothing deletable, orchestrator is the one THIN candidate (D-H9); D-H10 TDD shadow, D-H11 domain-modeling pairing added to the register. Prior-art: prior-art-evaluations.md#253 (same three-stack comparison; doc-only edit). * docs(arch-prep): slash-only planning skills evaluated as adoption candidates + 4 raw ideas Operator correction folded in: «no collision» ≠ «no value» — the user-invoked planning skills get per-skill adopt/adapt verdicts (wayfinder ADAPT strongest; to-tickets ADAPT mechanizable; to-spec one section; implement REJECT; triage two residues). New §4.6 carries four raw ideas for the design session: (1) the decision map as the multi-session layer over /arch — D4's register lifted to wayfinder shape, map-location sub-fork included; (2) kickoff Blocked-by edges with a pipeline-computed frontier; (3) seams-first Testing-seams slot in the spec template, unlocking the seams half of D-H2; (4) glossary SSOT as a term-ownership generated index — the CONTEXT.md-free adaptation that makes the grilling+domain-modeling pairing adoptable (D-H11 re-opened from defer). Register grows D-H12-D-H14. Prior-art: prior-art-evaluations.md#253 (same three-stack comparison; doc-only edit). * docs(arch): three-stack skill harmonization — design spec + continuation handoff Interview phase complete (frontier empty): P-1..P-6 operator premises, 15-area ownership map ratified (D-H1), decision register D-H0..D-H16 with falsifiers, mechanism set (prune script, CONTEXT.md rule+test, claim reorder, Blocked-by frontier, seams slot, aif plugin), probe register P1/P2a-c/P5-pending/P6, routed-work inventory for §3 exit routing. Awaiting §2 cold two-altitude review (this session's next step). * docs(arch): harmonization spec v2 — round-1 cold-review dispositions landed Both §2 seats returned REVISE (9 + 8 findings). All round-triggering findings repaired in place: §1 restated as two declared lanes (TD-F1); prune radius narrowed to 2 machine-globally-justified items per the operator's F7 answer + --check pre-push drift detector (TD-F2/F7); D-H17 completes the ownership map to all 11 model-invocable skills (TD-F3); setup run re-bucketed attended (TD-F4); D-H5 claim mechanics specified with real machinery + P4 restored (TD-F5, B-M1/M2); §5.6 non-target named (B-M3); /vitest transfer dissolved (B-M4); D-H7/D-H8 counter statuses corrected (B-M5); D-H13 adopts incumbent 'Depends on' spelling (B-M6). New: P-7 premise + D-H18 consumer-axis contour routed out via chip. Full dispositions: §9 v2 entry. * docs(arch): harmonization spec v3 + SSOT #253 counter arm + REJECT rows #254-257 Round-2 delta review (both seats REVISE; all round-1 closures confirmed): - --check channel corrected: owner:'maintainer' section in the pre-push.ts section registry (the file ships to consumers but maintainer sections never compose on a consumer layout, fail-closed) — .husky/pre-push is an exec dispatcher with no sections (convergent TD/B finding). - D-H16 build item DISSOLVED: aif container mounts the host ~/.claude/plugins read-only (docker-compose.override.yml), so the plugin is already visible in-container and the prune/--check cover it by construction (measured round-2). - 'counter armed' made true instead of re-worded: D-H7/D-H8 arm + observation No.0 appended to SSOT #253; REJECT rows #254-257 added (Matt implement, ADR dir, severity-less review model, total-sweep pruning). Spec SS8 item 5 DONE in-session. Dispositions: spec SS9 v3 entry. * docs(arch): harmonization contour GO — round-3 record + routed in-session edits Round 3 (targeted delta): both cold seats GO. Spec header → REVIEWED-GO; §9 round-3 entry (one TD MINOR accepted as recorded limit: container premise rests on untracked local docker-compose.override.yml — covered by D-H16 falsifier). Routed §8 item 2 small edits, per spec: - arch/SKILL.md §1: Testing seams slot added to the spec-template obligation (D-H14; seams-first adopted WITHOUT Matt's refactor placement) - ai-doc/SKILL.md: skill-authoring ownership note (standard=ours, process=SP writing-skills, writing-for-agents=REFERENCE) - rule-tests/SKILL.md: tautological-test anti-pattern REFERENCE note (D-H2 transfer (b)) * docs(arch): close harmonization contour handoff — full tail executed Review GO (3 rounds), exit routing done (3 chips + in-session edits), SSOT appends landed. Handoff retained as closure record; residue = operator actions (spec SS8 item 1) + chip-routed umbrellas. * docs(arch): consumer-axis satellite harmonization — design v1 + round-3 handoff Round-2 /arch contour (D-H18): interview closed, D-C1..D-C8 ratified with falsifiers; three-class collision model (factory CI / install-time census / informed consent); detect+declare+prescribe mechanism recorded. Cold review and exit routing DEFERRED behind the operator-mandated round-3 top-down creative re-examination (P-C3) — handoff written for the fresh session. * docs(arch): harmonization round 3 — registers amended, injected-context bindings land Round 3 (D-C8, operator-mandated P-C3) executed per the handoff's membrane phase order. Operator-axis spec v4: D-H15 SUPERSEDED — the prune apparatus (script / wizard / --check pre-push section / gate P5) dissolved, replaced by CLAUDE.md routing bindings (repo section + ~/.claude/CLAUDE.md machine-global half, written in-session with live operator approval) + meta-kickoff.template.md binding line (D-H10 fallback promoted to primary); D-H8 gains a frontmatter-neutering ladder step. Round-2 spec v2: D-C1 re-cut to the thin form (static census prose + known-pair presence check; inventory-join engine not built), D-C9 fourth-stack admission boundary added (knowledge-work trio stays on SSOT #235). Round-3 handoff closed with the continuation-state staleness correction; keen-shannon merged in (3ae6981) so both specs live on one branch. * docs(arch): P7 recorded — fresh-session bindings probe 2/2 vs P2 baseline Both P2-class triggers flip with the CLAUDE.md bindings in context (headless claude -p, fresh sessions reading the worktree CLAUDE.md from disk). Method finding recorded: in-session subagent probes are invalid for mid-session binding edits — subagents inherit the parent's session-start CLAUDE.md snapshot (measured via a failed in-session probe plus its diagnostic follow-up). * docs(arch): round-3 review R1 — both seats REVISE, dispositions landed Convergent BLOCKER fixed: the meta-kickoff.template.md binding line REMOVED — .claude/skills/pipeline/ ships to consumers via GETFF_SKILLS_ENV (setup.d/lib.sh:59) at the default env profile, so carrier #3 breached the operator-axis membrane while buying no coverage; its removal restores all 8 install fingerprints to the baseline blob. §5.1's «no mechanical channel at all» premise corrected (config layer only; frontmatter + a possible Skill-matched PreToolUse hook priced — P8 records the hook UNVERIFIED: guide claims no Skill matcher, live harness observation contradicts). P7 restated honestly (1 measured flip + 1 post-only confirmation). SSOT #253/ #257 got dated supersession notes (no prune ever executed). Five residual prune assertions re-cut. Consumer spec: population corrected — TWO shipped cc-plugin rows (superpowers + ast-grep, the latter disabled on the operator's own machine); presence check re-keyed on installed_plugins.json + enabledPlugins; D-C5/D-C6 aligned; class-2 own-skills half recorded as prose-only limit. ESCALATED to operator: ast-grep shipping fate (ESC-1) + the detection-wire fork (TD-M2/P8). ~/.claude/CLAUDE.md section relocated to file end (orphaned AIF bullet restored to its heading). * docs(arch): round-3 review R2 — residuals closed, operator answers landed Both R2 seats REVISE with a convergent root cause: R1 edited the surfaces findings argued FROM, not every surface repeating the claim. Closed: §1 premise re-cut to config-layer wording; §1 scope guard now names the pre-round-3 routed edits as verified degrade-safe REFERENCEs; sixth prune assertion re-cut (D-H16); handoff header unmerged label; D-H15 exclusivity hedge; consumer D-C1/§7 re-keyed on installed_plugins.json + enabledPlugins; both §8 inventories carry the escalations. Operator answers recorded live: ESC-1 → retro-census BOTH shipped rows, keep ast-grep on a clean census; P8 → VERIFY via the settings.json hand-off (§8 item 7). Review round cap (2 REVISE) reached — residual state surfaced in §9 instead of a third cold round. * docs(kickoffs): round-3 exit routing — two build umbrellas authored consumer-satellite-contract (thin form: retro-census of BOTH manifest rows per the answered ESC-1, D-C2 principle test, AGENTS.md.template section + parity line, install-registry-keyed presence check) and skill-harmonization-mechanisms (CONTEXT.md pointer-rule test, four-part claim machinery closing probe P4, Depends-on frontier). Both carry host-verify contracts and the PR-pause note: they become dispatchable only when the spec branch merges to staging. * docs(kickoffs): declare the effort-worthiness L0 rigor label on both round-3 kickoffs Principle 40 (`packages/core/principles/40-kickoff-rigor-label.test.ts`) requires every post-cutoff kickoff to carry a `Rigor label … L0 …` line with a legal value. Both round-3 kickoffs were authored without it and failed the gate at push time. - skill-harmonization-mechanisms → `build-and-verify`: all three surviving stages are factory-internal and reversible, each with a live RED/GREEN seam proof. - consumer-satellite-contract → `research-grade`: S3/S4 touch consumer-shipped surfaces (AGENTS.md.template, ./setup), which effort-worthiness §1 reserves for the research-grade contour. Prior-art: skipped — mechanical gate compliance on two doc files, no new capability * docs(arch): P8 CONFIRMED live — detection wire v0 declared; kickoff L0 labels The operator-registered log-only PreToolUse Skill hook fired on a forced model-invoked skill in a fresh headless session (JSON with tool_name=Skill + the skill name in tool_input). The guide-agent's 'skill loading bypasses the tool pipeline' claim is falsified — the P6 failure class again. Measured boundary: user-typed slash commands bypass the Skill tool (invisible to the wire, irrelevant: misroutes are model-invocations). TD-M2 closes — the log IS the v0 misroute detection wire feeding the D-H7 counter; spec §6 P8 + D-H8 + §8 item 7 updated. Both kickoffs gained effort-worthiness L0 rigor labels (parallel session's edit kept as-is). * test(install-sh): regenerate baselines for the three edited shipped skills The branch edits `.claude/skills/{ai-doc,arch,rule-tests}/SKILL.md` — all three are shipped artefacts, so their fingerprints move in every stack baseline that carries them. Captured with `SNAPSHOT_MODE=capture bash tests/install-sh/snapshot.sh`. Diff reviewed before committing (round-3 R1 precedent: an unreviewed template edit turned 8 fingerprints stale): exactly three payload paths changed hash — rule-tests (22 occurrences), arch (16), ai-doc (16) — and every fingerprint file is 1:1 on line count, so no payload entered or left any stack. Prior-art: skipped — snapshot regeneration after a shipped-file edit, no new capability --------- Co-authored-by: Test <test@example.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Wave 2 cleanup batch executing D-AuditC-1 R-1, R-3, R-4 (verify-only), R-6, R-7 + D6 + D-AuditC-6 from the 2026-05-16 strategic D-items dialogue (decisions.md, gitignored). All edits are documentation / metadata / cleanup; no rule introduction.
fix(agents): remove surviving broken self-testing-docs echo(acfdad8) — symmetric toc8fc153; line-41 echo referenced absentagents/references/path.docs(agents+CLAUDE): consumer-facing context note + ownership row per D-AuditC-6(0a1e3ad) — clarifiesdocs-auditor.md(scripts/audit-ai-docs.sh) andbest-practices-sidecar.md(.ai-factory/RULES.md) source-vs-consumer gap; adds CLAUDE.md Artifact Ownership Contract row.docs(skills): mark tool-bootstrapping Wave 5.3 as landed(c0efa1a) — tense correction (lands→landed).docs(research-patch): §17 think-time-gate erratum cross-link per D6(087bc7c) — preserves maintainer-flagged residual doubt (H2 vs H10 re-eval deferred to Phase 11+).Companion user-home edits (NOT in this PR): R-2 was a no-op (pre-cleaned), D-AuditC-3
mvexecuted manually by orchestrator after sandbox-blocked subagent attempt — see Test plan / ATTN.Test plan
--no-verify)ATTN (deferred — not actioned in this PR)
deps-hash-check.shis not registered in.claude/settings.jsonUserPromptSubmit hook (D-AuditC-1 R-3 observation). Scope-creep — surface for separate session.~/.claude/skills/ai-docs/SKILL.mdhad notoken-economyreferences at execution time (pre-cleaned state, 0 grep matches). No-op.mv ~/.claude/skills/git-user-info-ui-design.md→~/.claude/skills/projects/github-user-analytics/— subagent sandbox blocked the move; orchestrator executed manually after explicit maintainer confirmation.§1.7 Skipped: cleanup-only batch — no new rule, no new principle test, no new gate; CLAUDE.md row documents existing rule (D-AuditC-6), not a new one. Forward+Backward scope (per phase-research-coverage.md §1.7) protects rule-introduction drift, not mechanical maintenance of existing artefacts.
🤖 Generated with Claude Code (orchestrator session, Wave 2 of 2026-05-16 D-items execution)