feat: Wave 7 — hot-checks + harness-hooks + template test + §13.23 4th-layer - #29
Conversation
….27 + §13.28 + §13.8) O0 verdict NO end-to-end match — pre-commit/MegaLinter/Trunk don't unify harness-hooks with code+doc lints. Wave 7 scope: 5th lifecycle stage (harness-hook) + doc-linting + functional template test. §1.7 forward+backward applied; review session decides; no implementation in this commit. Prior-art: skipped — research patch, no new capability shipped (cites existing SSOT #1-#15; proposes new #16-#22 for sub-wave 7.x implementation commits)
…-1, MAJOR-2, MINOR-1, MINOR-2) MAJOR-1: §3 ATTN + REPORT ATTN — correct shipped-artefact paths (skills/, agents/ at repo root, not .claude/skills/, .claude/agents/); narrow Claude-first bias claim to .claude/skills/self-reflection/ + proposed harness-hook layer. MAJOR-2: §0 row + §1 row 1 + §8 entry #16 + §11 sub-wave 7.1 + §12.1 — lint-staged is NOT installed and root ESLint config does NOT exist; reframe as PRECONDITION for sub-wave 7.1 (not baseline). MINOR-1: cascade-table footnote — STILL ARMED includes deferred-with-armed-trigger. MINOR-2: §13.23 trigger #3 (pre-push surface widening) satisfied by Wave 7 sub-waves 7.1.b/7.1.c — explicit note in §0.1 + §6. Prior-art: skipped — text corrections to research patch, no new capability shipped
…§13.28+§13.8) Independent reviewer GO with confidence high. 6 PASS / 2 PARTIAL (write-time fixups, non-blocking) / 0 FAIL across 8 review dimensions per .claude/orchestrator-prompts/wave-7-hot-checks-joint-closure/review.md. §9 open decisions all closed: D1: ship harness-hook (5th layer) + §13.23 4th-layer design scheduled separately D2: DEFER Vale (FP risk on Russian+English mixed prose) D3: ship deterministic-only P1+P4+P6; LLM-judge probe DEFERRED D4: Claude Code only project-side; cross-editor stays WATCHLIST D5: §13.28 A+B both ship (sub-waves 7.4 + 7.2.b) D6: hand-roll .markdownlint.json (no Microsoft/Google preset) D7: harness-hook = SHOULD project-side / MAY consumer-side Sub-wave outline confirmed (7.1-7.5 per parent §11). Branch base note: Wave 6 not yet on main — operator merges Wave 6 first before Wave 7 PR opens. Prior-art: skipped — review verdict patch consolidating §9 open-decision closures, no new capability shipped
…hips in Wave 7 sub-wave 7.6 Decision 1 revised after user adversarial pushback («ты просто отложил потому что поленился делать?»). Adversarial check on load-bearing claim «§13.23 has 4 unresolved design problems unfit for Wave 7 scope» revealed initial deferral was false: Problem 1 (scope predicate): solvable via ## §-section diff content match Problem 2 (bootstrap chicken-and-egg): solvable via §1.7 Bootstrap: trailer Problem 3 (trailer-format interaction): NOT a problem — line-anchored parser unaffected Problem 4 (discipline-theatre): universal hook tradeoff, not §13.23-specific §13.23 4th-layer pre-push trailer check now ships in Wave 7 sub-wave 7.6 with full discipline cycle: 7.6.a research → 7.6.b independent review → 7.6.c implementation → 7.6.d closure (folded into 7.5.c). Wave 7 timeline: 6-9 days → 8-12 days (+2.5-3.5 days for 7.6). Sub-waves: 5 → 6. §1.7 enforcement ladder: 3 active + 1 deferred → 4 active layers post-Wave 7. §7 Revision note documents the failure mode for future review sessions: candidate anti-pattern #deferral-by-pattern-match-on-deferred-status — reviewer accepted «Status: deferred» + listed problems as load-bearing without running §1.4 adversarial check on each problem individually. Sample size 1/3 toward §3 distillation threshold. Prior-art: skipped — review verdict revision after adversarial-check correction, no new capability shipped
….23 per independent re-review Independent re-review (Agent task a618b8a650eb3b627, fresh session) returned REVISE on commit 4ede6ae (v2 revision). Findings: v2 overcorrected by claiming all 4 §13.23 problems tractable without verification — same cognitive shortcut as v1, opposite direction. Honest per-problem assessment: Problem 1 (scope predicate): UNCERTAIN — design sketch, pending 7.6.a Problem 2 (bootstrap): TRACTABLE — clean analogy to Prior-art: skipped escape-hatch Problem 3 (trailer-format): UNCERTAIN — claim rests on unverified parser mechanism assumption (git interpret-trailers vs grep line-anchored) Problem 4 (theatre risk): TRACTABLE — clean analogy to warn-only patterns Decision 1 routing: v1 (e75019a) = «defer» — accepted unsolvability without §1.4 check v2 (4ede6ae) = «CLOSED ship» — accepted tractability without §1.4 check v3 (this) = «CONDITIONALLY CLOSED» — 7.6 ships IFF 7.6.b GO on P1+P3 §7 documents symmetric anti-pattern pair: #deferral-by-pattern-match-on-deferred-status (v1) #tractability-by-pattern-match-on-claim-acceptance (v2) Generalised: #claim-acceptance-without-§1.4-on-the-claim-itself §8 adds operator-pre-approved fallback decision tree (F1-F7) for sub-wave 7.6.a/b outcomes — orchestrator does NOT escalate during execution if outcome matches pre-decided branch. Kickoff updates: §11 Decision 1 row reflects conditional closure; STOP-conditions add «7.6.b REVISE on P1/P3 → apply fallback §8»; mis-cite of independence rule corrected (now references Wave 7 review-prompt convention + §1.4 spirit, not §1.4 directly). Prior-art: skipped — review verdict v3 incorporating independent re-review findings, no new capability shipped
…ew findings (BLOCKER + 2 MAJOR + 2 MINOR) Second independent re-review (agent task afdc9688180da109f, fresh session, read-only) returned REVISE on commit 2be167f (v3). Plus self-check on load-bearing claims surfaced concrete evidence. Findings applied: BLOCKER (1, kickoff): CLOSEOUT block was unconditional «§13.23 closed by Wave 7». Orchestrator parses kickoff, not verdict patch — conditional was in the wrong document. Fixed: CLOSEOUT now has Path A (7.6 ships) / Path B (F6 fallback, §13.23 deferred to Wave 8) / Path C (operator escalation) branches. Sub-wave 7.6 header also carries explicit IFF condition. MAJOR-1 (Problem 4 overclaim): «warn-only-first-N-days is standard project pattern» asserted without verification. Self-check + agent caught it: self-application.md:60 documents Soft-warn concept in Decision matrix but NO husky-script implementation exists (grepped .husky/ + packages/core/ for warn-only / --warn — zero matches). Problem 4 downgraded from TRACTABLE-high → TRACTABLE-PENDING-CITE-medium. 7.6.a must design first- time soft-warn-in-husky implementation. MAJOR-2 (F6 cascade gap): if F6 fires (§13.23 defers), sub-wave 7.5.a row count changes 8 → 7. Not explicit in F6 description before. Fixed: F6 row now spells out 7.5.a row count change + 7.5.c closure note shape + 7.5.b SSOT entries still land (not wasted). MINOR-1 (Problem 2 hidden asymmetry): Prior-art: skipped is structurally triggered; §1.7 Bootstrap: relies on positional heuristic (rev-list count = 1) — fragile if hook commit not first on branch. Confidence downgraded HIGH → MEDIUM. 7.6.a must commit to either positional-with-hard-rule OR structurally-detectable subject-prefix. MINOR-2 (F4 arbitrary threshold): «3-day threshold» had no derivation from project metric. Replaced with operator-escalation trigger keyed to existing «>50% cost overrun» rule. Iteration pattern: v1 → v2 → v3 → v4 each surfaces smaller, less-prominent errors. Asymptotic-honesty convergence. Independent re-review at each step is what produces the convergence — not self-review. Prior-art: skipped — review verdict v4 incorporating second independent re-review findings, no new capability shipped
…cord hygiene) Third independent re-review (agent task a660e0834e5a89737, fresh session) returned GO with confidence medium-high on commit 5a5646a. v4 confirmed stable form — no BLOCKER, no MAJOR. Three MINOR cosmetic findings applied as record-hygiene v5: 1. §7 title renamed «TWO-STEP correction» → «Revision history — multi-step correction». Body always correctly described 4-step history (v1→v2→v3→v4); the heading was rotting and would mislead cold readers scanning headers. 2. §8 F4 escalation trigger now cites exact kickoff rule text verbatim («7.6.c implementation cost exceeds estimate by >50% (>2 days for ~1-2 day estimate)») instead of paraphrasing. Forward-tracing gap closed. 3. Kickoff sub-wave 7.6 header parallel-list extended «7.1/7.2/7.3» → «7.1/7.2/7.3/7.4». 7.4 doesn't conflict with 7.6 (different surfaces: Makefile vs .husky/pre-push); inconsistency with verdict §3 footnote resolved (footnote was already correct). Convergence verdict: three rounds of independent re-review converged on v4 as stable form. Errors changed CHARACTER across iterations (defer-overclaim → tractable-overclaim → calibrated-uncertainty), NOT just shrank in size — third agent flagged this nuance against the v4 commit message's «errors shrink» framing. The §7 «same cognitive shortcut» framing in the document is correct; the v4 commit-message «asymptotic-honesty» phrasing was slightly stronger than the document's own claim. v5 is the SHIP candidate. Implementation can proceed on orchestrator- kickoff basis per CLOSEOUT Path A (or fallback per F1-F7 in review §8). Prior-art: skipped — cosmetic record-hygiene revision; no new capability shipped, no decision logic changed
…endent-rounds heuristic (sample 1/3) Operator asked «когда останавливаться» — formalised the iteration-convergence signal that drove v4→v5 ship decision. Heuristic: «when N independent re-reviews converge on cosmetic/cite-fixing MINOR findings only with zero BLOCKER and zero MAJOR, ship the current revision». Proposed N=3 — mirrors AIF /aif-evolve aggregation threshold; empirically validated this iteration (Round 3 at cosmetic-only). Symmetric anti-pattern flagged: #never-ship-due-to-recursive-revision (asymptotic-honesty pursuit becomes process theatre when every round produces SOME finding regardless of marginal correctness gain). Sample size 1/3 — lives in §7 of this verdict per phase-research-coverage §3 «one patch per gap, append-only». Two more incidents needed for §3 distillation into rule §4 catalog (or new §5 «Stopping rules» section). (Also incorporates linter reformat of markdown tables — cosmetic widening of column padding; intentional per system reminder.) Prior-art: skipped — methodology distillation lives in §7 of review verdict, no new capability shipped
…aint analysis Operator-requested audit of all LLM touchpoints across project (current + planned + speculative). Constraint: «нет платного» — distinguish FREE (Claude Code subscription session-bound) vs PAID (CI / bare API). Key finding: current state is 100% «no-paid» compliant. - 11 current LLM touchpoints — ALL FREE (session-bound) - 12 CI jobs in audit-self.yml + discipline-self-check.yml + .husky/pre-push — ALL deterministic, zero LLM - Architectural invariant: «LLM-then-cache» pattern (LLM runs in session, writes deterministic artifact, CI reads artifact) §13.25 MCP discovery (operator's specific question): reuse AIF `/aif` for rules 1-4 (free, session-bound); rules 5-6 as in-session file-write. No standalone API-calling CLI. Documentation generation (operator's expected LLM use): ship as Claude Code skill (free, session-bound), NOT as CLI with API key. Wave 7 Decision 3 (LLM-judge probe): DEFER aligns perfectly with no-paid. CI gate deterministic-only; LLM-judge as optional manual session check if/when triggered. 5 open decisions surfaced (§9): D1: «no direct LLM in CI» as MUST row in Decision matrix D2: Consumer subscription assumption in INSTALL-FOR-AI.md D3: Gate 5 Opus subscription tier fallback policy D4: Doc generation as skill (recommended) vs CLI D5: Wave 5 constraint to session-only (no API-calling CLI) 2 SSOT proposal candidates (#23 LLM-then-cache pattern; #24 Claude Code hooks API) for future sub-wave commits. Prior-art: skipped — research patch documenting LLM usage audit, no new capability shipped
…tions (Decision 3 revised + D1/D2)
Operator selected Option B from 2026-05-11 LLM usage audit synthesis:
1. Decision 3 (§13.27 LLM-judge probe) REVISED — TWO-MODE shipment:
- CI gate (sub-wave 7.3.e) = deterministic only (P1+P4+P6)
- NEW local advisory skill (sub-wave 7.3.f) = Claude Code session-bound
LLM-judge for paraphrase/cue-placement checks. FREE within subscription.
Audit confirmed deterministic-CI + session-LLM is the architectural sweet
spot for «no-paid» constraint. Operator's «documentation for AI validated
by AI» philosophical point honored without triggering §13.11 Gate 5
cost-tracking infrastructure prematurely.
2. D1 from audit — «no direct LLM API calls in CI» added as MUST meta-rule
row in self-application.md §3 Decision matrix at sub-wave 7.5.a.
Locks the «LLM-then-cache» architectural invariant against future drift.
Decision matrix row count: 8 → 9 (Path A) / 7 → 8 (Path B F6 fallback).
3. D2 from audit — INSTALL-FOR-AI.md update at sub-wave 7.2.d now covers
BOTH editor-coupling acknowledgement (Decision 4) AND «requires Claude
Code subscription for AI-assisted workflows» note (D2).
D3 (Opus tier fallback), D4 (doc generation as skill speculative §13.X),
D5 (Wave 5 HARD CONSTRAINT) — DEFERRED per operator Option B choice.
Live as armed triggers; handled in respective Wave 5 / Wave 8+ scopes.
Wave 7 timeline: 8-12 days → 9-13 days (+1 day for B additions).
Sub-waves remain 6 (7.1-7.6); sub-wave 7.3 gains 7.3.f.
Kickoff updated in parallel (operator-side, gitignored):
- ARGUMENTS estimated days
- §11 Decision 3 row revised
- Sub-wave 7.2.d expanded with D2
- Sub-wave 7.3 gains 7.3.f
- Sub-wave 7.5.a 8 → 9 rows
- CLOSEOUT Path A 8 → 9; Path B 7 → 8
Prior-art: skipped — review verdict v6 incorporating Option B audit-driven additions; no new capability shipped at this commit
…p injection Add .claude/settings.json with UserPromptSubmit hook entry and .claude/hooks/inject-session-bootstrap.sh (13 LOC). The hook prints a session-bootstrap digest (goal phrase + 3 invariants + Step-0 reading order + pointer to .claude/session-bootstrap.md) to stdout; Claude Code injects stdout into the user-prompt context automatically. Closes Wave 6 D-1: session-bootstrap injection via harness-hook channel. DECISIONS: - .claude/settings.json was ABSENT before this commit. - New-file LOC: settings.json 14 + hook script 13 = 27 total — under 80 threshold. - Capability commit: NO (total new-file LOC <80; no new packages/ subdirectory). Prior-art: skipped — config-layer hook addition (UserPromptSubmit channel) extending settings.json with session-bootstrap injection digest; no new packages/ subdirectory or ≥80 LOC file. Closes Wave 6 D-1 (session-bootstrap injection via harness-hook).
- Add markdownlint-cli2 ^0.17.0 as workspace root devDependency. - Create .markdownlint.json (root, auto-discovered): 6 hand-rolled structural rules (MD001/MD003/MD007/MD009/MD034/MD040); no preset. MD013 line-length disabled via default:false (long prose lines are legitimate). - Create packages/lint-config/ (workspace member @rules-as-tests/lint-config): mirrors root config in markdownlint.json + documents rule set in README.md (99 LOC, Authoritative-for header per doc-authority-hierarchy rule §3). - Extend .husky/pre-commit with markdownlint-cli2 check on staged *.md files (ACMR diff-filter; bash-driven, no lint-staged). Rule set decisions: MD001 — heading hierarchy; skipped levels hide doc structure drift. MD003 atx — project uses # prefix exclusively; setext style rejected. MD007 indent:2 — 4-space indent creates parser rendering differences. MD009 — trailing spaces are invisible noise; no intentional <br> in prose. MD034 — bare URLs skip link text; [text](url) required for context. MD040 — code blocks without language tag degrade syntax highlighting. MD033/MD013 disabled — prose-style FP risk; 500-line limit covers length. Capability commit: YES (new explicit dep markdownlint-cli2 in package.json; new file packages/lint-config/README.md ≥80 LOC under packages/). Prior-art: prior-art-evaluations.md#17 (markdownlint-cli2, verdict transition DEFER → ADOPT after operator-side discipline rule §13.27 closure).
…apper (§13.28 B) Add .claude/hooks/validate-prompt.sh (34 LOC) and PostToolUse Edit|Write hook entry in .claude/settings.json. The hook fires on every Edit/Write tool call, filters to .claude/orchestrator-prompts/**/*.md paths via shell-glob match, then invokes validate-batch-spec.ts via node_modules/.bin/tsx. Exits 0 silently on pass or unmatched path; non-zero + diagnostic on red. Gracefully skips (exit 0 + stderr warn) if jq or tsx unavailable. exit 2 from validate-batch-spec.ts (gh CLI unavailable) is treated as soft-skip (exit 0) to avoid blocking tool calls for missing tooling. DECISIONS: - Invocation form: node_modules/.bin/tsx (resolved REPO_ROOT-relative) rather than bare npx tsx, to avoid auto-install attempts. - Path globbing: shell pattern [[ ... == *".claude/orchestrator-prompts/"*".md" ]] (not jq-aware) — Claude Code PostToolUse passes file_path in tool_input; no need for jq-level glob, plain shell match suffices. - Capability commit: NO (wrapper ≤50 LOC + settings.json hook entry; no new packages/ subdirectory or ≥80 LOC file). Prior-art: skipped — config-layer wrapper script ≤50 LOC + settings.json hook entry invoking validate-batch-spec.ts on .claude/orchestrator-prompts/**/*.md PostToolUse events. Closes §13.28 part B (operator-side discipline via Claude Code hook channel). No new packages/ subdirectory or ≥80 LOC file.
Extend .husky/pre-push with §8 lychee offline link check on changed *.md files (git diff origin/main..HEAD --diff-filter=ACMR). Checks broken relative links and anchors without network access (--offline --no-progress). Design decisions: - Offline mode: no network calls — deterministic and fast; remote URL validity is out of scope for a pre-push hook (belongs in CI link-check workflow). - Absence fallback: if lychee not in PATH, print one-line warning + install path and exit 0. Rationale: operator-discovery responsibility; hard-fail would block all pushes on fresh checkouts before lychee is installed. Install path documented inline: cargo install lychee OR brew install lychee. - File selector: git diff origin/main..HEAD --diff-filter=ACMR | grep -E '.md$'. Skips cleanly when no changed .md files exist (empty CHANGED_MD check). - Word-split on $CHANGED_MD (unquoted): safe for this repo's filename conventions (no spaces in doc paths); avoids xargs -r portability issue on macOS BSD xargs. Capability commit: NO (config-only; no new packages/ subdirectory, no new file ≥80 LOC; modification to existing .husky/pre-push — status=M not status=A). Prior-art: skipped — pre-push hook addition (config layer) wiring lychee --offline check on changed Markdown; no new packages/ subdirectory or ≥80 LOC file.
…a+7.3.b)
Add packages/core/audit-self/template-render.audit.ts — vitest .audit.ts
harness for §13.27 functional template test. Commit 1 (skeleton): install.sh
wiring in beforeAll, P1/P4/P6 probes stubbed as .todo pending commit 2.
Vitest discovers the file via new audit-self/**/*.audit.ts include in
packages/core/vitest.config.ts.
Add skeleton fixtures under packages/core/audit-self/fixtures/:
- ts-server/package.json: {"name":"ts-server-fixture","version":"0.0.0","private":true}
- react-next/package.json: declares next@^15 + react@^19 (declaration-only, no install)
- {ts-server,react-next}/src/.gitkeep: empty src/ placeholder
- fixtures/README.md: folder-level authority header per doc-authority-hierarchy.md §5
Decision log:
- Skeleton fixture shape: 2 stacks × 1 package.json + src/.gitkeep = 6 files
- README.md in fixtures/ — folder authority pattern applied (§5)
- afterAll: rm -rf tmpdir via node:fs/promises rm — no repo pollution
- beforeAll timeout: 90s (install.sh is file-copy only; alarm if slower in CI)
Prior-art: prior-art-evaluations.md#22 (Cookiecutter/Copier, verdict ADOPT VOCABULARY — vocabulary-only, no Python dep, no Copier runtime adopted; functional template test pattern named «template-render audit» per Copier's «answers.yml replay» analog).
…b-wave 7.1.c)
Extract REQUIRED_HEADER_DOCS + check logic from 09-doc-authority-hierarchy.test.ts
into a new sibling module 09-doc-authority-hierarchy.ts with stable exported API:
export const REQUIRED_HEADER_DOCS: readonly string[]
export const EXEMPT_PATTERNS: readonly RegExp[]
export function isExempt(relPath: string): boolean
export function hasAuthorityHeader(content: string): boolean
export function checkDocsHaveAuthorityHeader(
paths: string[], repoRoot?: string
): { ok: boolean; violations: Array<{ path: string; reason: string }> }
CLI shim 09-doc-authority-hierarchy.bin.ts:
- Reads paths from process.argv.slice(2)
- Calls checkDocsHaveAuthorityHeader(paths)
- Prints violations to stderr; exits non-zero on !ok
- Invokable via: npx tsx packages/core/principles/09-doc-authority-hierarchy.bin.ts <paths>
- Follows project's existing tsx-based bin pattern (matches rules-as-tests-detect etc.)
Test file (09-doc-authority-hierarchy.test.ts):
- Now imports from new module instead of inlining logic
- No behavioural change to existing 10 tests; 4 new API smoke tests added (14 total)
- New tests: ok:true for valid path, violation for missing file, EXEMPT_PATTERNS export,
mutation: exemption suppresses violation in changed-files mode
Stable API consumed downstream by sub-waves 7.2.c (PostToolUse) and 7.4
(make validate-prompts). repoRoot parameter defaults to process.cwd() so
both hook contexts (repo root) and test context (REPO_ROOT) work without
additional configuration.
Capability commit: NO (refactor of existing principle 09 to expose changed-files
mode for pre-push / harness-hook integration; no new packages/ subdirectory
or ≥80 LOC new file; status=M on test + status=A on extraction module).
Prior-art: skipped — refactor of existing principle 09 to expose changed-files mode for pre-push / harness-hook integration (consumed by sub-waves 7.2.c and 7.4); no new capability shipped.
…y header check Add .claude/hooks/check-doc-authority.sh (29 LOC) and second PostToolUse Edit|Write hook entry in .claude/settings.json. The hook reads the edited file path from stdin JSON, converts it to a repo-root-relative path (REQUIRED_HEADER_DOCS API expects relative), then invokes the Batch A CLI shim via node_modules/.bin/tsx. The CLI filters to REQUIRED_HEADER_DOCS — exits 0 for unrecognised paths. DECISIONS: - Batch A's 09-doc-authority-hierarchy.bin.ts WAS visible at smoke-test time (landed in the working tree at 10:15 on 2026-05-11). Test path: README.md → "OK 1 path(s) checked". Batch A is not yet on a merged commit; recorded as known runtime dependency in case merge ordering changes. - Path conversion: ABS_PATH stripped of REPO_ROOT prefix to obtain relative path, since checkDocsHaveAuthorityHeader() resolves relative to process.cwd() (= REPO_ROOT when called by npx tsx from repo root). - Capability commit: NO (≤30 LOC wrapper + settings.json hook entry; reuses Batch A capability; no new packages/ subdirectory). Prior-art: skipped — config-layer hook entry + ≤30 LOC wrapper invoking principle 09 CLI from sub-wave 7.1.c (Batch A). No new capability shipped; reuses changed-files mode from 7.1.c.
…+7.3.d)
Wire install.sh invocation in template-render.audit.ts beforeAll: spawnSync
against ts-server + react-next fixture skeletons, each in a fresh os.tmpdir()
workdir. Implements three deterministic probes per Decision 3 (no LLM in CI):
P1 — goal-phrase rendered (AGENTS.md.template lines 13/57):
Synonym list: «enforced by lint, tests, CI» | «without a measurable check».
Source: packages/core/templates/shared/AGENTS.md.template (documented paraphrase
variants that actually appear in rendered output; verbatim README phrase absent).
P4 — Authoritative-for header (duplicated from principle 09; authority header RE).
Checks rendered AGENTS.md carries "> **Authoritative for:**" header.
TODO: import from principle 09 shared module once Batch A lands.
P6 — taxonomy fidelity: every framework skills/ entry exists in consumer
.claude/skills/ after install. Catches silent rename drift.
Each probe includes an anti-tautology mutation test (bare content → probe fails),
mirroring the negative-test discipline in audit-ai-docs.test.sh.
P2/P3/P5 (LLM-judge) intentionally NOT shipped per Decision 3 + «no-paid» /
«LLM-then-cache» invariant. They live in .claude/skills/template-audit/SKILL.md.
Decision log:
- Synonym list source: AGENTS.md.template lines 13+57 (verbatim, not paraphrased)
- REQUIRED_HEADER_DOCS: duplicated inline (AUTHORITY_HEADER_RE) with TODO to
consolidate post-Batch-A-merge; full rendered-path mapping deferred to that
consolidation to avoid duplicating a 30-entry list
- install.sh invocation: spawnSync, 60s per-stack timeout
Prior-art: skipped — extends template-render.audit.ts from commit 1 with install.sh wiring + P1/P4/P6 deterministic probes. No new packages/ subdirectory or ≥80 LOC new file. P2/P3/P5 LLM-judge probes intentionally NOT shipped in CI per Decision 3 (two-mode: deterministic CI / advisory local).
…arch Prior-art: skipped — research patch for §13.23 4th-layer closure, no new capability shipped.
…ave 7.1.d)
Wave 6 D-3 SHIP-B independent track folded as cross-ref in sub-wave 7.1.d.
D-3 probe added to packages/core/audit-self/audit-ai-docs.sh:
- Checks that canonical goal phrase from README.md §Goal appears verbatim
(or via synonym) in enumerated downstream goal-bearing docs.
- Canonical phrase: "AI agents can't silently bypass undocumented conventions"
- Synonym: "AI cannot silently bypass what fails CI"
- Downstream docs checked (explicit enumeration, not glob):
.claude/session-bootstrap.md — operational restatement; must carry goal
CLAUDE.md — AI-tooling conventions; must carry goal pointer
- Probe ID: D3; selectable via --only=D3 flag.
- Smoke-run result: PASS (both downstream docs carry the canonical phrase).
Context7 sweep (≥3 phrasings, 2026-05-11; no production analog surfaced):
1. «doc-vs-doc parity / drift detection» → Drift SDK (DeFi), anomaly
detection resources — no doc-parity tool surfaced.
2. «AI documentation paraphrase drift template rendered output» →
Documentation.AI (documentation platform), Mistral/Google Document AI
(OCR/extraction) — different capability; no parity checker.
3. «documentation template consistency downstream parity» → Doc Detective
(tests docs against product behaviour, not doc-vs-doc phrase parity),
NgDoc (Angular docs generator) — different domain.
Verdict: BUILD — hand-rolled probe, no production analog.
SSOT entry #16 appended to docs/meta-factory/prior-art-evaluations.md:
Candidate: goal-phrase parity check (context7 sweep 2026-05-11)
Verdict: BUILD
Trigger to revisit: dedicated doc-parity tool surfaces; OR 7.3.d P1 probe
generalises to cover this case; re-evaluate at 7.5.b.
Also fixes MD040 violations in co-staged files (caught by new pre-commit hook):
- docs/meta-factory/prior-art-evaluations.md:57 — entry template fenced block
- docs/meta-factory/self-application.md:70 — ASCII path diagram
- INSTALL-FOR-AI.md:45,104,130,181 — copy-paste prompt / output blocks
All fixed by adding "text" language tag. Hook working as designed.
Capability commit: NO (status=M on existing audit-ai-docs.sh — not status=A;
no new packages/ subdirectory or ≥80 LOC file per Wave 6 MAJOR-1 correction).
Prior-art: skipped — extension of existing audit-ai-docs.sh with D-3 goal-phrase parity probe; no new packages/ subdirectory or ≥80 LOC file. Context7 sweep performed: (1) doc-vs-doc parity / drift detection; (2) AI documentation paraphrase drift template rendered output; (3) documentation template consistency downstream parity — no production analog surfaced. SSOT entry #16 appended with verdict BUILD.
Add .github/workflows/framework-self-template-render.yml: new CI job that runs npm --prefix packages/core run test:template-render (vitest .audit.ts harness from commits 1+2). Also adds test:template-render script to packages/core/package.json (mirrors test:principles pattern). PARTIAL dim 4 write-time fixup: Local runtime measurement: 1.3s on Apple M-series / Node 22 / macOS 25.4. Target ≤60s; alarm if CI runtime >120s on baseline ubuntu-latest runner. If CI runtime exceeds target consistently — surface as research-patch on fixture-size vs probe-coverage trade-off. Deterministic-only gate: job env contains NO AI-vendor secrets. No ANTHROPIC_API_KEY, no curl/fetch to Anthropic or OpenAI. Enforces «no-paid» / «LLM-then-cache» invariant. LLM-judge probes (P2/P3/P5) are session-bound in template-audit skill. Decision log: - New workflow file (not appended to audit-self.yml): cleaner review scope - Job pin: actions/checkout v4.2.2 + actions/setup-node v4.4.0 (mirrors audit-self.yml) - Node 20 (mirrors audit-self.yml baseline) Prior-art: skipped — CI workflow addition wiring framework-self-template-render job invoking template-render.audit.ts from commits 1+2. Deterministic-only per Decision 3 + «no-paid» / «LLM-then-cache» invariant. No new packages/ subdirectory or ≥80 LOC file.
Add .claude/skills/template-audit/SKILL.md — 42-line session-bound advisory
skill for template render QA beyond the deterministic CI gate.
Two-step procedure:
Step 1: runs npm --prefix packages/core run test:template-render (deterministic
CI-equivalent P1/P4/P6 probes from commit 1+2). Stop on failure.
Step 2: asks the current Claude Code session (no API call, FREE under subscription)
to semantic-check rendered content for P2/P3/P5 advisory findings:
P2 — paraphrase fidelity (goal-phrase framing not drifted to «guidelines»)
P3 — cue placement (session-bootstrap cue in first 10 lines of AGENTS.md)
P5 — synonym coverage (P1 list still matches current template phrasing)
NOT a CLI, NOT a CI gate, NO API key per «LLM-then-cache» invariant. Session-bound
under Claude Code subscription. Findings are advisory (not blocking).
Auto-trigger keywords: template, audit, render, generated docs, AGENTS.md,
paraphrase, cue placement, local advisory, template-render, audit-template.
Promotion trigger: P2/P3/P5 → CI gate when deterministic PASS rate <80% over
30 days OR consumer goal-phrase miss-detection report fires (per Decision 3
re-evaluation, open-questions.md §13.27).
Decision log:
- Skill location: .claude/skills/ (same as self-reflection, the existing pattern)
- Slash command /audit-template: deferred — .claude/commands/ does not exist in
this repo; no convention to follow; skill auto-trigger keywords are sufficient
- Trigger-keyword list final: 10 keywords covering all entry surfaces
Prior-art: skipped — local advisory skill ≤50 LOC, session-bound under Claude Code subscription, no CLI, no API key per «LLM-then-cache» invariant. Promotion to CI gated by Decision 3 re-evaluation trigger (deterministic-only PASS rate <80% over 30 days OR consumer goal-phrase miss-detection report).
Add verification row to INSTALL-FOR-AI.md checklist: "Harness hooks active (Claude Code only) | jq .hooks .claude/settings.json | UserPromptSubmit + PostToolUse entries present" Companion to 7.2.a/b/c hook files already committed. Consumer can verify harness-hook layer wired correctly post-install with a single jq call. Prior-art: skipped — doc-only addition to INSTALL-FOR-AI.md (one checklist row); no new capability shipped. Editor-coupling section and self-application.md preview row landed in 2b0a505 as co-staged with 7.1.d. Full §3 matrix expansion lands in sub-wave 7.5.a.
…e 7.4.a) Appends `validate-prompts` to existing Makefile. Walks .claude/orchestrator-prompts/**/*.md (excluding README.md) and runs validate-batch-spec.ts on each file; exits non-zero on first failure. Closes §13.28 part A (operator-side Makefile discipline layer). Prior-art: skipped — operator-side make-target wrapping existing validate-batch-spec.ts validator on .claude/orchestrator-prompts/**/*.md; no new packages/ subdirectory or ≥80 LOC file. Closes §13.28 part A (Makefile discipline layer).
…ew verdicts Prior-art: skipped — review verdict patch for §13.23 4th-layer, no new capability shipped
…b-wave 7.4.b) Adds README.md with Authoritative-for header per .claude/rules/doc-authority-hierarchy.md §5 (folder-level authority pattern). Declares append-only status and validate-prompts validation channel. Updates .gitignore: replaces directory-level ignore with file-glob pattern (.claude/orchestrator-prompts/*) + explicit negation for README.md so the folder README is tracked while individual prompt files stay gitignored. Subdirectory prompt files remain ignored (subdirectory entries are ignored by the * rule, so git does not descend into them). Prior-art: skipped — doc-only README applying folder-level authority pattern per .claude/rules/doc-authority-hierarchy.md §5; no new capability shipped.
Independent reviewer audit of sub-waves 7.1 (Batch A) + 7.2 (Batch B) + 7.3 (Batch C) + 7.6.a (Batch F.a) against orchestrator-kickoff.md spec. Domain scores: 1. Commit hygiene WARN (2b0a505 co-staging; SSOT #16 pre-land) 2. Hard constraints PASS 3. 7.1 hot-check prims PASS 4. 7.2 harness-hook layer WARN (settings.json UserPromptSubmit nesting) 5. 7.3 template test WARN (P1 synonyms are methodology phrases) 6. 7.6.a research patch PASS 7. Cross-batch consistency PASS (typecheck clean) 8. Capability compliance PASS (80ef1d9 + f528586 Prior-art trailers OK) Overall: GO WITH NOTES — 3 WARN items (Issues 2+4 must verify pre-PR; Issue 5 carry-forward to 7.5.d). 3 NOTE items for orchestrator. ATTN: 7.6.b already on branch with GO [F1] verdict — 7.6.c is unblocked. Prior-art: skipped — review verdict patch for Wave 7 Round 1, no new capability shipped.
…7 review) The PostToolUse hook for Authoritative-for header check was firing on every Edit/Write — including .ts/.json files that are not authority-bearing docs. Root cause: `09-doc-authority-hierarchy.bin.ts` passed argv paths to `checkDocsHaveAuthorityHeader` without filtering. Only `EXEMPT_PATTERNS` were skipped inside the function; any non-doc file outside that pattern produced a violation, surfacing as `FAIL …: missing header` on stderr. Fix: filter argv to `REQUIRED_HEADER_DOCS ∪ EXEMPT_PATTERNS` BEFORE calling the check. Empty filtered list → silent exit 0. Three new tests cover (a) non-required path silent, (b) required-without-header FAIL, (c) exempt silent. Bonus: OK-message rewritten so exempt-only paths no longer report «all have header» (review §2 m5). Discovered by cold-start reviewer (2026-05-11) via empirical hook smoke-test with non-doc path; previous round audit (Domain 4) declared «hook scripts all correct» without running this probe — gap captured in reviewer-checklist amendment (Commit 3 of this batch). Prior-art: skipped — fix M1 from Wave 7 Round 1 review (non-doc paths leak through CLI shim); no new packages/ subdirectory, no new file ≥80 LOC, no behaviour-extension beyond regression fix.
…#22 Cookiecutter/Copier (Wave 7 review) Two prior Wave 7 commits cite SSOT entries that did not yet exist: 80ef1d9 (sub-wave 7.1.a markdownlint, capability=YES) → #17 f528586 (sub-wave 7.3.a/b template-render harness, capability=YES) → #22 The kickoff (§Sub-wave 7.5.b) planned SSOT entries #17–#22 to land batch-style at Wave closeout, but CLAUDE.md's build-vs-reuse invariant mandates: «add a new SSOT entry … in the same commit as the capability artifact.» CLAUDE.md wins; the kickoff plan was wrong for #17 and #22 (which are positively cited from prior commits). For #18–#21 (lychee, vale, Claude-Code-hooks, cross-editor scripts), the corresponding sub-wave commits used `Prior-art: skipped — …` escape-hatch trailers rather than positive citations, so they are NOT in violation — the kickoff plan for those four remains intact and they land at 7.5.b alongside their verdict transitions. Resolution: forward-add (not amend / not force-push) — the cited trailers in 80ef1d9 and f528586 now point to existing rows. CLAUDE.md «prefer create new commit than amending» is honoured. The deviation was caught by Wave 7 Round 1 cold-start review (2026-05-11); prior round audit Domain 8 reported «PASS» by verifying trailer presence without checking SSOT existence — gap captured in reviewer-checklist amendment (Commit 3 of this batch). Prior-art: skipped — forward-fix M2 from Wave 7 Round 1 review (SSOT entries #17 + #22 added in same atomic commit per CLAUDE.md same-commit rule; supersedes the deferred-to-7.5.b plan in kickoff for these two entries only; entries #18-#21 land at 7.5.b as planned because their corresponding commits did not cite missing IDs); no new capability shipped.
…test additions Two corrections to the Round 1 audit (47dc9ed, docs/meta-factory/research-patches/2026-05-11-wave-7-round1-review.md): - Issue 4 (settings.json UserPromptSubmit nesting): DEBUNKED — hook fires empirically (this very session's UserPromptSubmit digest is proof). - Domain 8 (Capability compliance) + Domain 4 (Hook scripts): both re-classified PASS → WARN — method gaps (trailer-presence-only vs SSOT existence; no empirical hook smoke-test). Forward-fixed by Commits 1+2 of this batch. The phase-research-coverage rule (.claude/rules/phase-research-coverage.md) gets two new §1 checklist items so future review sessions catch these gaps proactively: §1.8 Hook surface smoke-test (PostToolUse / pre-push) with path outside hook's intended filter → expect exit 0 §1.9 SSOT citation existence-check: `grep` for #N in prior-art- evaluations.md, not just trailer presence Both items are derived from gaps that cold-start review surfaced 2026-05-11 but the prior round audit missed. Rationale for landing in the rule file (vs a new reviewer skill): the file already owns §1 adversarial-check discipline; the two smoke-tests are concrete executions of that discipline for hook-touching / capability-commit changes. Prior-art: skipped — doc-only errata + skill smoke-test addition; no new capability, no new packages/ subdirectory, no new file ≥80 LOC.
…+ §13.23 self-review Updates self-reflection skill (.claude/skills/self-reflection/SKILL.md) by adding the §1.7 enforcement layers table: 3 active + 1 deferred → 4 active. The 4th layer (pre-push trailer check) shipped in Commit 1 of this batch. Ships paired self-review patch at docs/meta-factory/research-patches/2026-05-11-§13.23-4th-layer-self-review.md per §13.23 promotion path #4. The patch walks the new layer through §1.7 forward+backward checks applied to its own motivating gap (local-push- bypasses-CI + discipline-theatre risk). §13.23 status update in open-questions.md folds into sub-wave 7.5.c per kickoff §«Sub-wave 7.6.d» (atomic with §13.27 + §13.28 closures); NOT modified in this commit. §1.7 self-application: this commit touches .claude/skills/self-reflection/ SKILL.md AND adds new ## §1.7 enforcement layers section heading. The C4 predicate fires on this commit. The §1.7 check shipped in Commit 1 of this batch (one commit earlier in the push range) and is in warn-only calibration mode — B1 bootstrap trailer below satisfies §1.7 for this commit. Prior-art: skipped — doc-only update to self-reflection skill description + self-review patch in research-patches/; no new packages/ subdirectory, no new file ≥80 LOC. §1.7 Bootstrap: updating skill ladder from 3 active + 1 deferred → 4 active layers and shipping self-review patch immediately after introducing the §1.7 enforcement layer — §1.7 check itself is in 30-day calibration (warn-only) at this commit; B1 exemption per 7.6.a research §4 + 7.6.b §2 Problem 2 verdict (mirrors Commit 1 bootstrap exemption; both commits are in the same push, same bootstrap context).
…ers + 1 meta-rule) Wave 7 closeout matrix update — appends 9 rows to self-application.md §3 Decision matrix Phase 1. Each row cites §13.8 4-criteria gate verdict (failure-cost / local-cost / detectability / lifecycle stage). 1. markdownlint-cli2 structural — MUST (pre-commit; failure-cost MEDIUM, local-cost LOW, detectability MEDIUM, lifecycle 1) 2. lychee offline link check — SHOULD (pre-push; MEDIUM/LOW/HIGH/2) 3. Harness-hook UserPromptSubmit+PostToolUse — SHOULD/MAY (real-time; HIGH/LOW/MEDIUM/5; full row replacing preview) 4. Functional template-render audit — MUST in CI (deterministic; HIGH/MEDIUM/HIGH/3) 5. Local advisory template-audit skill — MAY (session-bound; LOW/ZERO/HIGH/1) 6. Operator-side make validate-prompts — SHOULD (Makefile; MEDIUM/LOW/HIGH/1) 7. §13.23 4th-layer pre-push trailer check — SHOULD→MUST (warn-only 30-day calibration; HIGH/LOW/MEDIUM/2) 8. (meta-rule) No direct LLM API calls in CI — MUST (architectural; HIGH/ZERO/LOW/3) — codifies «no-paid» invariant per LLM-usage audit 9. Goal-phrase parity audit — SHOULD (pre-push scoped; MEDIUM/LOW/MEDIUM/2) Row 3 (harness-hook) replaces the line-62 preview row with full rationale + promote-to-MUST trigger. §1.7 forward+backward check applied. Forward: 9 rows comply with R1-R20, principles 01-09, Prior-art trailer, doc-authority — no contradictions surfaced. Backward sweep: each row mapped against existing matrix rows 53-61; complementary additions or expansions, no conflicts. §13.8 «Decision matrix expansion mechanism» dogfooded (9 rows added in one atomic commit per the §13.8 4-criteria gate format) — updated in 7.5.c. Prior-art: skipped — doc-only matrix expansion enumerating shipped Wave 7 layers, no new capability or new file ≥80 LOC.
… verdict transitions) Append 4 SSOT entries to docs/meta-factory/prior-art-evaluations.md per Wave 7 closeout (§3 «Adding a new entry» step 1): #18 Vale — DEFER (Russian+EN mixed prose FP risk; promote when 2nd doc- prose incident markdownlint cannot catch) #19 lychee — ADOPT (transition from ADOPT WHEN TRIGGERED; shipped 7.1.b) #20 Claude Code hooks — ADOPT (UserPromptSubmit + PostToolUse; shipped 7.2.a/b/c; transition from ADOPT WHEN TRIGGERED) #21 Cursor/Cline/Continue.dev reference scripts — WATCHLIST (Claude Code only project-side per Decision 4; cross-editor parity deferred) NOTE on numbering (entries #16, #17, #22, #23-#26 landed in earlier commits per CLAUDE.md same-commit-as-capability rule): #16 2b0a505 (sub-wave 7.1.d D-3 BUILD verdict) #17, #22 Batch G «feat(prior-art): M2 …» (forward-fix from Round 1 cold-start review M2 finding) #23-#26 Batch F.c Commit 1 «feat(hooks): Wave 7 7.6.c — §1.7 trailer check + SSOT #23-#26» (atomic SSOT landing per CLAUDE.md; precedent 7.1.d #16; cited #23 + #26 positive trailer) The kickoff §«Sub-wave 7.5.b» plan originally targeted all #16-#22 at 7.5.b — corrected via Round 1 review M2 finding adjudication + Round 2/3 user patch confirming atomic #23-#26 landing in F.c Commit 1. Prior-art: skipped — appending parent-research SSOT entries at Wave 7 closeout per §3 «Adding a new entry» step 1; no new capability, no new packages/ subdirectory.
…update §13.8 Atomic closure of three armed §13.x triggers and update of one running trigger per Wave 7 closeout: §13.23 (4th-layer pre-push trailer check) — CLOSED. Implementation shipped in Batch F.c Commit 1 (.husky/pre-push section 9). Status: "deferred 2026-05-09 → closed by Wave 7 sub-wave 7.6, 2026-05-11". §13.27 (functional template test) — CLOSED. Implementation shipped in sub-wave 7.3 (deterministic CI gate framework-self-template-render) + 7.3.f (local advisory skill template-audit). LLM-judge re-evaluation trigger added per Decision 3: promote when deterministic PASS rate <80% over 30 consecutive days OR Claude Code subscription expands to CI compute. §13.28 (operator-side discipline) — CLOSED. Implementation shipped in sub-wave 7.4.a (Makefile target validate-prompts) + 7.2.b (PostToolUse harness-hook running validate-batch-spec.ts on .claude/orchestrator- prompts/**/*.md). Both A and B per Decision 5. §13.8 (Decision-matrix expansion mechanism) — UPDATED (status remains open as the running mechanism trigger). Wave 7 sub-wave 7.5.a dogfooded the mechanism — 9 rows added citing 4-criteria gate verdict per row. Also fixes 2 pre-existing MD040 (fenced code without language) violations at lines 18 and 192 (unlabeled code blocks now use ```text). Prior-art: skipped — open-questions status updates per Wave 7 closure; no new capability, no new packages/ subdirectory.
…tim README synonym
Two artefacts closing Wave 7 self-review:
1. New research patch 2026-05-11-wave-7-decision-matrix-self-review.md —
§1.7 backward sweep applied to the 9 new Decision matrix rows landed
in sub-wave 7.5.a (Commit 1 of this batch). Walks 9 rows × R1-R20 +
principles 01-09; surfaces complementary / extension / no-overlap
verdict per pair. Non-trivial pairs called out:
- row 7 (§1.7 trailer) × P08 (prior-art cited): different keys +
different hooks, coexistence proven by Batch F.c
- row 4 (template-render audit) × P09 (doc-authority): static vs
install-time rendering — complementary layers
- row 1 (markdownlint structural) × 500-line limit: additive checks,
different enforcement mechanisms, no duplication
All 9 rows verdict: complementary to existing R1-R20 + P01-P09.
Ratifies sweep cadence: any Decision matrix expansion ≥3 rows triggers
a self-review patch (mirrors §13.8 4-criteria gate dogfooding).
2. m1 fix per Round 1 cold-start review (2026-05-11):
packages/core/audit-self/template-render.audit.ts — GOAL_PHRASE_SYNONYMS
now includes verbatim README#why-this-exists phrase as third synonym.
Drift detection now covers three independent expressions: two paraphrase
variants currently in templates + verbatim README for forward compat.
All 9 template-render vitest tests still PASS.
Subject prefix `docs(research-patches):` triggers D3 §1.7 allow-list
bypass (per .husky/pre-push section 9 v1) — no §1.7 trailer required
because the patch is doc-only and the m1 fix is an audit-config tweak.
Prior-art: skipped — self-review patch + audit synonym addition per Wave 7 closeout; no new capability, no new file ≥80 LOC.
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
|
Warning Review the following alerts detected in dependencies. According to your organization's Security Policy, it is recommended to resolve "Warn" alerts. Learn more about Socket for GitHub.
|
…(CI fix) Removes duplicative DECISIONS log section (lines 492-497) — the same decisions are documented in body sections §0 (context7 sweep), §5 (synthesis), §6 (SSOT), and §Final (bootstrap path). Net delta: −7 lines. Post-trim: 498 lines, comfortably under the audit-self 500-line invariant. Triggered by audit-self / Mechanical checks failure on PR #29. Prior-art: skipped — CI fix removing duplicative content from research patch; no new capability, no new packages/ subdirectory.
…verlap analysis Adds three entries from the aif-handoff-overlap-analysis mandate (Stage 3, Phase A). Entries reflect verdicts from Stage 1 research + Stage 2 review: - #27 DEFER: AIF Handoff HANDOFF_MODE env-var fork (C1×P4 ORT — different problem classes; DEFER pending non-interactive orchestration pipeline) - #28 DEFER: AIF Handoff paused:true/false semantic (S1×P5 ORT — different granularities; DEFER pending wave state machine formalisation Phase 11+) - #29 ADAPT: AIF first-line plan annotation pattern (A1×P6 OVL; ADAPT as <!-- scope:§N --> for research patches + automated trigger sweeps §1.6) Phase B (ADAPT implementation: annotations + CI gate) follows in separate commit, citing entry #29 as Prior-art trailer. Prior-art: skipped — adding SSOT entries to prior-art-evaluations.md; no new capability artifact, no new file under packages/, no new package.json dep.
…-> (SSOT #29) Implements ADAPT verdict from aif-handoff-overlap-analysis mandate (Stage 3). Adapts AIF handoff:task first-line plan annotation pattern (SSOT #29) for research patches as <!-- scope:<slug> --> machine-parseable first-line marker. Changes: - 20 research patches: added <!-- scope:<slug> --> as first line per canonical mapping (§13.21, §13.23, §13.25, wave-7, phase-8.8, methodology, aif-handoff-mandate) - packages/core/principles/10-research-patch-annotation.test.ts: companion principle test with anti-tautology mutation check (Vitest, 5 tests green) - .github/workflows/audit-self.yml: CI gate verifying annotation presence on all research patches (except README.md), added to mechanical job - 2026-05-10-wave-6-review-verdicts.md, 2026-05-10-wave-7-review-verdicts.md: fix pre-existing MD040 bare fences (language=text) uncovered by pre-commit hook Enables automated §1.6 trigger sweep: grep "scope:§13.X" across patches instead of prose-level pattern matching on filenames. §1.7 backward check: all 20 existing patches annotated atomically with CI gate in this commit (complete sweep, not §1.5 floor). Prior-art: prior-art-evaluations.md#29 (AIF handoff:task first-line plan annotation ADAPT — <!-- scope:slug --> for research patches, automated §1.6 sweep support, verdict ADAPT 2026-05-11).
…l disposition (#162) Completes the DN-4 sweep: - CONTRIBUTING.md «Working in a git worktree» (#29 worktree_node_modules_symlink) - phase-research-coverage.md §1.13 — AI-doc research source priority (#30) - tracker: #29/#30 → CODIFIED; #18/#19 → DEFERRED (reference-facts); #21 → DEFERRED (thesis-level); #24 → OUT-OF-REPO (global orchestrator skill) DN-4 now fully swept: 11/15 codified, 4 dispositioned with reasons, zero bare-PENDING. §1.7 self-review at research-patches/2026-05-22-dn4-round3-tail-codification.md. Prior-art: skipped — prose discipline codification + tracker disposition of existing memory conventions, no new capability or dependency
… Implementer-equivalent only) value-add audit (#276) R-phase patch for Sub-wave C of the aif-handoff-as-runtime-bridge umbrella. Evaluates Variant C (kickoff §3 lines 124-145): aif-handoff as Implementer- equivalent only, bypass Planner+Reviewer cycle, thin CLI wrapper for kickoff dispatch + kanban status tracking. Verdict: REJECT (BFR-default §1 ladder). Rationale: - The kickoff-framed "aif-handoff exec --kickoff <path>" CLI does not exist in lee-to/aif-handoff (DeepWiki probes 1+5, 2026-05-29). - No first-class Implementer-only mode; skipReview:true bypasses Reviewer but Planner is mandatory unless accept_existing_plan with on-disk PLAN.md (same disk coupling SW-A flagged for Variant A). - BEFORE/AFTER maintainer-action count: 25% literal / 0% cognitive reduction (T-AIF-BRIDGE-C table §4) — below kickoff §8 STOP 30% threshold → verdict "Variant C value-add insufficient". - Pure-tracker pattern (paused:true + autoMode:false + manual state-machine transitions) IS shipped but adds zero automation beyond UI tracking; Docker+SQLite infra unjustified. Cites: - SW-A merged PR #268 (Variant A REFERENCE, 28% match, 3 ADOPT-blockers) - SW-B merged PR #267 (Variant B REJECT, ~5% match, no dir-watch capability) - PR #269 follow-up (mechanical corrections, no verdict changes) - DN-1=B-constrained input consumed in criterion 5 (mooted for Variant C which bypasses aif-handoff Reviewer entirely) - Gate-4 admission re-sweep: PR #127+#128 touch packages/runtime/ only (no MCP/coordinator drift in 30-day window) 5 distinct DeepWiki probes + 2 WebSearches + cross-ref to SW-A/SW-B/PR #269 = 19+ evidence channels (T1 floor exceeded 3.8x). §1.7 forward+backward + §self-application + T-trap walk per ai-laziness-traps.md §3. Single output file under docs/meta-factory/research-patches/. No code, skill, agent, install.sh, or .claude/rules/ modifications. ### §1.7 Forward-check applied build-first-reuse-default.md §1 verdict ladder applied; BFR §3 6-layer search performed (SSOT rows #27/#28/#29/#30/#43/#44/#46/#67/#80 reviewed at prior-art-evaluations.md:95-148; DeepWiki >=5 probes; WebSearch >=2 phrasings; own-stack sweep at .claude/skills/meta-orchestrator/SKILL.md:441 anti-scope + :404+:429 SP requesting-code-review). no-paid-llm-in-ci.md §1 enforced (all evidence via subscription-bundled DeepWiki/WebSearch + free gh CLI + bash). reviewer-discipline.md §2 respected (DN-1=B-constrained consumed as fact, not re-litigated; verdict is research finding against §8 STOP, not strategy choice). ai-laziness-traps.md §3 active T-traps applied (T1, T3, T7, T11, T12, T13, T15, T16, T17, T19, T20, T-AIF-BRIDGE-C MANDATORY BEFORE/AFTER table at patch §4). Evidence: see patch §8 file:line citations. ### §1.7 Backward-check applied SSOT #27/#28/#67 receive additive notes (additive-only; no verdict changes). Original DEFER/DEFER/REJECT rationales reviewed at prior-art-evaluations.md: 95, 96, 135 — consistent with Sub-wave C findings (reinforce existing classifications, do not re-litigate). No .claude/rules/* modified; no .claude/skills/* modified; no agents/* modified; no packages/* modified; no install.sh modified; no kickoff.md modified. Single output file in docs/meta-factory/research-patches/. Scope strictly bounded to Variant C; SW-A/SW-B/SW-B2/SW-D out of scope. T15 self-application confirmed in patch §10. Memory not written (Sub-wave D synthesis is the natural codification surface). Evidence: see patch §9 file:line citations.
…+ population report (#1117) * docs(umbrella-donemd-backfill): Commit A — done.md for 12 CLOSED-VERIFIED candidates Stage 1 of umbrella-donemd-backfill. Closes 12 stale no-done.md dirs where closure is proven by either parent-umbrella's done.md naming the per-stage PR (3 cases) or parent-umbrella's final PR closing the meta-launch runtime state (4 cases) or direct branch-name PR match where the branch IS the slug (5 cases). Per-candidate evidence (full matrix + population enumeration in report.md landed in Commit C): agnosticism-remediation-t-a/b/c (parent done.md names per-stage PRs): - T-A #570 — CC-qualify skill auto-activation + doc-claims probe - T-B #575 — portable AGENTS.md rule index + rules-autoload coverage probe - T-C #577 — agnosticism-conformance principle test slot 21 Parent (agnosticism-remediation) closed by #589 (2026-06-16). *-meta-launch dirs whose parent umbrella is closed (meta-launch runtime state included in parent's final PR): - consumer-upgrade-path-meta-launch ← parent #615 - egress-host-push-default-meta-launch ← parent #760 - generator-require-composite-tier-meta-launch ← parent #708 - hook-test-suite-rot-meta-launch ← parent #608 Branch-name PR matches (branch IS slug; no kickoff.md to verify final-stage trap T-UDB-C against; dir is runtime batch state for the named work): - wave-5-trio-followup ← PR #37 (chore/wave-5-trio-followup, 2026-05-11) - wave-7-hot-checks-joint-closure ← PR #29 (wave-7-hot-checks-joint-closure, 2026-05-11) - wave-8-retro-h8-gate5-13-31 ← PR #48 (chore/wave-8-retro-h8-gate5-13-31, 2026-05-12) - phase-9-implementation ← PR #18 (docs/phase-9-implementation, 2026-05-09) - audit-fixes-2026-05 ← PR #1 (feat/audit-fixes-2026-05, 2026-05-07) Structural note (load-bearing, surfaced in report §1): all 12 target dirs lack kickoff.md (untracked + nothing on disk). priority-score.sh skips such dirs at line 144 (`if [[ ! -f "${kickoff}" ]]; then … continue`) — so writing done.md here is documentation, not a priority-score.sh panel-state change. The 65 no-done dirs with kickoff.md are all excluded by either the explicit live-program list or the 45-day live-signal rule, so Stage 1 yields zero priority-score.sh panel changes (correctly: there are no actionable closures in the live-program-with-kickoff population). Prior-art: skipped — refactor only, no new capability * docs(umbrella-donemd-backfill): Commit C — report.md + .gitignore exception Stage 1 report for umbrella-donemd-backfill. Adds the population enumeration, per-verdict counts, the full UNCLEAR table (22 rows with per-row signals), and an acceptance-gate self-check. Load-bearing structural finding in §1: 35 of 100 no-done.md dirs are runtime-only coordination slots (no kickoff.md on disk, no git history). priority-score.sh already filters these out at line 144 (kickoff-missing -> continue) BEFORE the Layer C3 done.md check at lines 222-231 runs, so the C3 layer never reaches them whether done.md exists or not. The 12 done.md files from Commit A add documentation value (visible closure marker) but do not change /pipeline panel output for those 12 -- they cover the 12 of 65 real-umbrella no-done.md dirs where closure was provable via parent-umbrella or branch-name PR evidence. Per-verdict counts (per §2 of the report): - CLOSED-VERIFIED: 12 (done.md landed in Commit A) - SUPERSEDED: 0 (none of the 34 candidates has kickoff.md -> S3 superseded-marker check is inapplicable) - OBSOLETE-CANDIDATE: 0 (no per-dir visibly-overtaken evidence possible without kickoff context) - UNCLEAR: 22 (weak/ambiguous signals; first-class outcome per T14) .gitignore: extends the orchestrator-prompts exception list with report.md parallel to done.md/kickoff.md. Same durable-committed class -- the report is an audit artefact, not runtime state. Acceptance gate (§7 of the report): all 6 checks pass. priority-score.sh runs clean (the new done.md files land in already-skipped dirs + the 12 real-umbrella closures). ZERO done.md files under exclusion-list dirs. UNCLEAR section present. Prior-art: skipped -- prose-only doc commit, no new capability. * docs(umbrella-donemd-backfill): expand report §7 with full gate evidence Task 6 acceptance gate is complete. Updates §7 to record all 6 gate checks explicitly + triages the 5 pre-existing done.md files in exclusion-list dirs (multi-model-pipeline-pilot, launch-preannounce-track, python-delivery-v0, generation-live-delivery, meta-orchestrator-prior-art) as pre-existing state authored by other maintainers' past PRs, not leaks from this umbrella. `git log --author=Yhooi2 --since=2026-07-23` confirms zero of my commits today touched any of those 5 files. All 6 acceptance-gate checks pass: 1. priority-score.sh runs clean (exit 0, 367 lines, no parse breakage) 2. ZERO NEW done.md under exclusion-list dirs (5 pre-existing are not this umbrella's work) 3. Every done.md corresponds to a report row with cited PR + merge SHA 4. UNCLEAR section present with 22 rows (T14 compliance) 5. Population enumeration present with exact command + 100-dir count 6. Schema check passes on all 12 new done.md files Prior-art: skipped -- prose-only doc edit, no new capability.
Summary
Wave 7 joint closure ships 6 sub-waves of hot-checks, harness-hooks, functional template test, and §13.23 4th-layer pre-push enforcement. Closes 4 armed open-question triggers in one branch via atomic dogfooding of the §13.8 Decision-matrix expansion mechanism.
Path A confirmed — §13.23 4th-layer (
§1.7pre-push trailer check) ships warn-only under 30-day calibration. Total: 36 commits / 42 files (+3091 / -149).Sub-wave breakdown
validate-prompt+ PostToolUsecheck-doc-authorityframework-self-template-renderCI gate (deterministic P1/P4/P6 probes) + local advisorytemplate-auditskillmake validate-promptsMakefile target + folder-level authority README.husky/pre-push §9§1.7 trailer check (warn-only 30-day calibration) + SSOT #23-#26 atomic landingOpen-questions closed by Wave 7
framework-self-template-rendershipped (7.3)make validate-prompts) and B (PostToolUse) shipped (7.4 + 7.2.b)§1.7 enforcement ladder: 3 active + 1 deferred → 4 active layers (rule + skill + CI workflow + pre-push trailer check).
Decision matrix expansion (sub-wave 7.5.a, +9 rows)
Each row cites §13.8 4-criteria gate (failure-cost / local-cost / detectability / lifecycle-stage):
make validate-prompts— SHOULD§1.7 forward+backward sweep applied to all 9 rows — no contradictions with existing R1-R20 + principles 01-09 + Prior-art trailer discipline. Full sweep documented in
docs/meta-factory/research-patches/2026-05-11-wave-7-decision-matrix-self-review.md.SSOT register state (cumulative)
All entries land same-commit-as-capability per CLAUDE.md build-vs-reuse rule (precedent: 7.1.d/#16).
Round 1 cold-start review findings (2026-05-11)
Independent reviewer surfaced 2 MAJOR + 9 MINOR + 1 DEBUNKED. Closed in Round 3 «Batch G» fix bundle:
check-doc-authority.shleaked FAIL noise on non-doc paths (CLI shim no filter); fixed viaselectRequiredPaths()extraction in09-doc-authority-hierarchy.bin.ts(commit eab7c71)..claude/rules/phase-research-coverage.md§1.8 (hook smoke-test) + §1.9 (SSOT citation existence-check) added (commit 4ec69e5) so future review sessions catch these gaps.Recursive self-application (§1.7 sweep)
§1.7 Forward-check applied
The Wave 7 shipment introduces new disciplines (§1.7 4th-layer pre-push trailer enforcement, §1.8 hook surface smoke-test, §1.9 SSOT citation existence-check, 9 new Decision matrix rows including the «No LLM API in CI» architectural meta-rule). Forward-check verifies each new discipline complies with currently-active layers:
npm typecheck --workspacesPASS across 3 workspaces.npm --prefix packages/core run test:principlesPASS 51/51. Principle 08 (Prior-art cited) PASS — every capability commit (80ef1d9, f528586, 91f5f1d, 5c0d32e, 2e43874, 38dc50a, 71a7f00, 29e62d8, 5afabad, 8982fde, 2b0a505, 951d7f7, eab7c71) carries validPrior-art:trailer (positive citation or escape-hatch with ≥20-char rationale).> **Authoritative for:** …header where required. Principle 09 PASS — REQUIRED_HEADER_DOCS list compliance verified, EXEMPT_PATTERNS preserved.§1.7 Bootstrap:trailer per B1 exemption; F.c Commit 2 (e5ada1a, SKILL.md ladder 3→4) also carries B1 (C4 fires on SKILL.md edit per F.b §4 fixup Phase 4: Stack Detector v1 (read+write AIF bridge, CI gate, reviewer cycle complete) #4). All other commits either fall outside §1.7 scope predicate (file-glob mismatch or D3docs(research-patches):allow-list bypass) or land before the §1.7 layer ships (warn-only fires on pre-enforcement commits 4ec69e5 + 91f5f1d, exits 0).§1.7 Backward-check applied
Backward sweep applies new rules to existing artefacts under their scope. For Wave 7:
docs(research-patches):/chore(snapshot-regen):/chore(prior-art-update):prefixes); meta-test: positive (B1 trailer with ≥20-char rationale satisfies §1.7), mutation (removing B1 detection breaks bootstrap path).docs/meta-factory/research-patches/2026-05-11-wave-7-decision-matrix-self-review.md(caa65dc, 7.5.d self-review patch). Backward-check ratifies sweep cadence: any Decision matrix expansion ≥3 rows triggers a self-review patch.GOAL_PHRASE_SYNONYMSintemplate-render.audit.ts:42-45— extends coverage without removing existing methodology phrases;npm --prefix packages/core run test:template-renderstill PASS 9/9.Known limitations / expected behaviour
origin/main..HEADrange).S17_WARN_ONLY=falseper TODO comment in.husky/pre-push. 50-commit adversarial test FP rate = 0% (well below 5% threshold for permanent warn-only fallback)..claude/rules/phase-research-coverage.mdAuthoritative-for header (header still says «6-item» while body now has 9 items §1.1-§1.9) — cosmetic, pre-existing drift (was 6 vs 7 before Wave 7), deferred to follow-up cleanup or Wave 8.Hand-off to Wave 5 implementation
Wave 5 (project-aware tool bootstrapping) implementation has NOT yet started — Wave 7 first per 2026-05-10 sequence reversal. Wave 7 surfaces infrastructure Wave 5 should adopt:
framework-self-template-renderCI gate.make validate-prompts— Wave 5 orchestrator prompts get advisory validation channel.Test plan
npm --prefix packages/core run test:principles: 51/51 PASSnpm --prefix packages/core run test:template-render: 9/9 PASSnpm run typecheck --workspaces: 3 workspaces PASSbash -n .husky/pre-push: syntax OKmake validate-prompts: all files passed (self-reflexive on Wave 7 prompts themselves)make self-audit: PASS (5/5 audit probes)§1.7 Bootstrap:) on F.c commits grep-resolvablepackage.json/.tspaths now exit 0 silent through PostToolUse hook