v4.2.0 — the spec diff reaches Build
Eleven new scripts, seven new references, and four path bugs that all
produced the same failure: a confident false negative.
Stage 3 now receives the spec diff, not just the spec. A superseded
requirement keeps its original text, so a removal was invisible to a run
reading only the current file — the dropped behaviour survived in the code.
Stage 5 derives its coverage denominator from the spec rather than from the
check's own declaration, so a check can no longer shrink its own scope and
pass. Team-level settings are honoured only from the committed config; a
local override of one is ignored with a warning, and disabling every check
is a load error rather than a clean run over nothing.
New:
- spec-diff.sh the delta handed to Build, with the no-baseline cases kept apart
- drift-reverse.sh code citing a requirement that no longer stands
- validate-spec.py EARS and id permanence, executable: 61 codes, ERROR/WARN
- signals.sh/score.sh evidence and judgment split, judgment keyed to an evidence hash
- req-trailer.sh Productizer-Req: provenance, orphans, COV_ coverage
- init.sh all of Stage 0 in one command, with an interview fallback
- graduate.sh repeated corrections into durable guidance
- retrieval-budget.sh / ab-harness.sh / stage-snapshot.sh the measurement suite
- usage-audit.sh our own scripts' error rate and option contract
Onboarding no longer dead-ends. The survey forks to a weaker evidence tier
instead of refusing, and gained probes for CLI surface, public API, config
keys, CI jobs and skill inventories. Two of four real repositories tested
could not be imported before this, including this one.
Fixed, all four the same class — a relative path anchored to the wrong
directory, reported as an absence:
- spec-diff.sh reported a tracked spec "absent" from any subdirectory
- run-checks.sh resolved ROOT, the default config and --changed against
three different wrong places; it also wrote a nested shadow .claude tree - stage-status.sh counted with
grep -c || echo 0, which emits two zeros - import-survey.sh --help printed bash's
cdbuiltin help
Docs corrected against the code: GUIDE.md and README.md still named
.claude/sdlc/ and the lettered stages; SKILL.md claimed six stages while
defining nine, and stage 5 had no heading at all. The gitignore rule in
SKILL.md took its verdict from git check-ignore -v, which exits 0 on a
negation and so refused exactly the repos that had already fixed theirs.
Measured, and it is worse than we said: the symbolic cross-check scores
1.00/0.70 on the 21 cases we wrote first and 0.67/0.12 on a harder 26-case
corpus. The figure tracked the corpus, not the checker. B3 records both.