v4.7.0 — thirteen pending items opened
B11-B23 record what this session found and did not fix: the agent halts on a
contradiction without asking for the ruling; there is no cross-session memory
at all; a check has no rendering after a human overrides it; nothing asserts
one living spec per product; eleven requirements have no acceptance row; the
runner format ships with no executor; the measurement suite has never met a
real model. B9 is closed - --strict is clean - and B10 carries the graded
result.
Making a never-triggered check amber, which shipped in 4.5.0, was most of a
good idea and one bad one. A check scoped always that did not run is broken.
A check scoped by paths that did not match is not applicable to this change,
and counting it turned the tile amber on every commit that touched no shell
script - which is how an amber signal stops meaning anything, three commits
after being built to mean something.
run-checks.sh now records trigger_scope, and the view reads it. Found by
falsifying rather than by looking: the first attempt failed because the view's
check-row normaliser built a fixed dict and dropped the new key, so every check
read as scoped and the real repo happened to render correctly for the wrong
reason. Two arithmetic errors in the same neighbourhood were caught the same
way - "3 check(s), all passing · 1 not applicable" of three checks is two
claims that cannot both be true, and the partial denominator counted a check
that never ran among those that had.
Four states, falsified: scoped miss alone is green and says what did not apply;
an always check that did not run is amber; both together are amber and count
each separately; a real failure is red.