A correctness release for the review surface. Two of the fixes below mean
that on v0.8.0 every design was approved against a picture that was not the whole
picture, and the gate meant to catch a dropped call could report clean without
having judged anything.
Added
TESTING.md— a value→test map, enforced by a test. Seven properties, each
stated by its failure, with every file undersrc/test/javamapped to a
property or a declared bucket.CoverageGovernanceTestchecks it
both ways: every cited test resolves, is tracked by git and carries no
@Tag("eval"); every test file is accounted for. Guarding against "a mechanism
existed and did not run" with a review convention would have been the same kind
of object as the mechanisms that failed, so the map executes.- CI, for the first time.
.github/workflows/test.ymlruns./gradlew test
on push and pull request. It installs Node, becauseFrontendChainTestaborts
when node is absent and would have skipped 18 tests while reporting green — the
defect this repo keeps hitting, arriving through the tool meant to catch it. The
workflow fails when the skip count moves off its expected value. - The frontend harness can reach a success path. It loads the real
diagram-svg.js, lets a case stubfetch(the default still refuses the
network), and records DOM writes so a test can assert what a panel rendered.
Until nowdrawModel's success path was unreachable from any test, which is
why the defect below survived.
Fixed
- The reviewer was never told what derivation could not account for. Stage A
records a capture gap for a branch, a loop, a chained or static receiver, a
source that will not parse — and the derive response carried every one of them
insideslice.captureGapsall along. The wizard readdata.sliceModeland
discardeddata.slice, so the Before panel drew a partial flow and said nothing
about the rest:slice-act1shows 3 arrows for an 8-call method, and this
repo's ownderiveByPathderives 8+ gaps. The panel now names each gap above
the diagram, and states complete capture rather than leaving silence to be
interpreted. Honesty that never reaches the reader is not honesty. No backend
change was needed.CaptureGapDisclosureTest. - The dropped-call gate counted calls the design does not make.
proposedCallSet()resolved the sequence itself and omitted the caller check
resolveSteps()applies, so a step whose caller did not resolve emitted no
arrow while the gate still treated the call as proposed — the design lost a
call silently and the gate reported nothing. A false release, in the gate built
to catch exactly that. It now reads the shared selection.ReviewGateTest. - Editing a frontend test driver did not re-run the tests. Four Java tests
shell out tonode src/test/js/*.js, andCoverageGovernanceTestparses
TESTING.mdandWHY.md— none of which Gradle can see, because they are read
from the project directory rather than the classpath. A content change left
:testUP-TO-DATE and reused the previous results, so editing a claim switched
off the check that polices claims.build.gradlenow declares all of them as
task inputs. - The reviewer could not see the entry interaction or the return type.
emitPlantUmlemitted[*] -> SUTand[*] <-- SUT;renderSequenceDiagram
had no system-caller branch, sofindParticipant('__system_caller__')returned
undefined and its guard dropped both steps. Step 3's after panel is that
renderer, so every design was signed off against a picture missing two
arrows the file carried anyway. A third copy of the resolution indexed fragment
brackets and had to agree with both. All three now read oneresolveSteps()—
the same shapeSliceRendereralready used server-side, where one
pumlArrows()sits behindrenderPumlandrenderModel"so the drawn diagram
and the puml text never disagree".DesignAgreementTestcompares the drawn
labels against the emitted arrows; the emitted text is byte-identical. - The dropped-call gate failed open, and stuck that way.
renderReviewBefore
cleared the gate before deriving and recomputed only on success, so a derive
error left it reporting "no dropped calls" after judging nothing. The failure
was cached, so every later visit to Step 3 replayed the banner and skipped the
recompute too. A failed derive now blocks sign-off with its own message;
greenfield, which genuinely has no baseline, still passes. The same gap existed
while the derive was merely in flight —renderReviewBeforereturns
synchronously and the fetch settles later — so that window blocks too. Failures
are no longer cached: caching one turned the fix into a trap that locked
sign-off for the whole session on a single transient error, with nothing able to
release it.ReviewGateTest.
Changed
- The claim set is corrected. An adversarial read of
WHY.mdfound seven
defects. The framing sentence claimed each claim forces the next; it does not —
claim 5's mechanism belongs to claim 1, claim 7's to claim 3, and claim 6 stands
alone. Claim 5 said the reviewer understands the code, which contradicts
"leaf content is sampled"; it now says structure. Claim 6 said "drift cannot
return" without qualification; design/code drift dies, intent drift moves to the
PR record and ages there. Claim 7 said "check all of it"; it checks that every
designed interaction has a test. Claim 1's PRs are self-approved and now say so.
Claim 3's naive counterfactual is hand-authored and now says so, together with
the chain-1 run that confirmed it. Claim 4 drops NPath for change cost.
Two things were missing and are added: a non-claim about the single shape DisC
imposes, and an answer to the objection the file never faced — why review at
all, when you can make failure cheap and regenerate? - Derivation now proves fidelity as well as stability.
CallGraphDeriverSoundnessTest
proves every call in an entry body is captured as an arrow or disclosed as a
capture gap. Stage A now records a gap for a source that will not parse, a file
it could not read, two provided types sharing a simple name, and a call Stage E
cannot emit an arrow for — a static receiver or an unprovided callee type. Each
one used to pass silently, so a slice built on an incomplete world looked exactly
as confident as a complete one, and REGEN would have overwritten the orchestrator
without whatever it missed.captureComplete()had promised this all along. - Analyze no longer spends a model call on plugin validation. The
deterministic half of what the plugin's Step 1 would refuse now runs locally
and instantly as part of the design lint; the plugin's own check runs once,
at Generate, where a refusal is about to cost a real generation. One fewer
model call on every Analyze, and the wizard reports most problems immediately
rather than after a round trip. - The refusal retry is gone. It fed the plugin's grammar refusal back to
the sequencer, which only produces the ordered call list — nothing it emits
could fix a sealed family with one permit or an unpinned decision. It cost a
model call and could not have worked. (The data-flow retry, which sends
argument problems to the component that authors arguments, is untouched.) - Validation asks a capable model, not Haiku. Step 1 is deterministic
rule-checking written as prose and executed by a model, so the verdict varies
with the model: the same design was refused by Haiku and accepted by Sonnet on
the same day. Cheapness mattered when this ran on every Analyze; it runs once
now, so it can afford to be right.
Fixed
- A refused generation said nothing at all. Moving the plugin's check to the
Generate button left the panel that displays its verdict back on Step 2, inside
a section that isdisplay: nonewhile you are on Step 4. Clicking Generate on
a design the plugin refuses showed "Checking the design…", returned to
"Generate", and stopped — no message, no generation, no explanation. The panel
now lives in the step it fires from, andWizardMarkupTestfails if it moves
again. - Hand-authored designs were refused for a field the wizard never sets. The
contract checks require every entity to name an owning participant, a rule
written for the analyzer's output. Nothing the wizard builds carries it —
neither the "+ Entity" button nor the entities inferred from method signatures
— so an ordinary design reported "entity X has no ownedBy", which blocked
sign-off and wasted a model call on a retry that could not have fixed it. The
owner is now derived from the participant whose signature introduced the type. - Grammar problems were being sent to the sequencer. The lint endpoint
returned data-flow and contract findings in one list, so a "sealed family has
one permit" reached the retry that exists for data-flow problems, wrapped in a
prompt about call arguments — the same misrouting removed from the validator
path, arriving through a different door. The two families are now reported
separately and only data-flow findings drive the retry. - The sign-off gate labelled contract problems as data-flow problems. They
rendered under "this design uses values that nothing in the flow produces",
behind a checkbox reading "these values arrive some other way — proceed".
Neither sentence was true of them, and ticking it waved through something the
plugin would hard-refuse minutes later. Contract findings now get their own
block, their own wording, and an acknowledgement that says what it does. - The contract checks ran before the design was assembled. They fired inside
the sequencer step, which completes before the wizard picks the SUT and adds
the entry interaction — so they judged a half-built model and reported "sut is
missing or blank" every time. They now run at sign-off, on a finished design. - Validation could silently stop validating. The verdict parser only read
the plugin's{"ok": …}envelope when it was the very first thing in the
output, so any run where the model narrated before answering fell through to
the "unparseable → pass" branch. The gate reported success without judging
anything. The envelope is now found anywhere in the output, and a refusal is
recognised first so one that quotes the envelope cannot read as a pass. - The design validator now sees the decision tables.
--validate-only
received the.pumlalone, so Step 1 judged the design as if every leaf were
unspecified: a decision table could contradict the diagram, or specify a call
the design never makes, and validation still passed — the reviewer was told the
design was fine while the generator went on to read a different one. Sidecars
are now staged beside the design being validated. (Long-standing; tracked as
"P4".) - A Stage-D guard that could never fire. The "SUT must not be added" rule
required the change to be filed underelement == "participant", but the
differ files SUT rewiring under"arrow", so no real delta could ever trip it.
Now keyed on the operation and name, which is the actual invariant.
Added
- Tests that prove the claims in WHY.md, all free of model calls:
DerivationStabilityTest(repeated derivation is byte-identical; a comment
moves nothing; one added call moves exactly one arrow),
DesignDeltaValidatorTest(Stage D had no direct tests despite owning the rule
that prevents a wholesale overwrite from deleting uncaptured code),
DesignReceiptTest(the orchestrator has zero branches while the naive
counterfactual for the same ticket has three; the expectedverify()count is
computable from the design alone). DerivationStabilityForeignRepoTest— the same stability properties
against any repository, opt-in and free:
./gradlew test --tests '*ForeignRepo*' -Ddisc.stability.repo=<path> -Ddisc.stability.entry=Class#method. Verified on upstreamspring-petclinic.- WHY.md — the seven-claim argument for DisC, each claim with its
mechanism, the test that proves it, and an explicit list of what DisC does
not claim.