Repository navigation
TestGuard v0.6.0
Everything the field asked for. Two independent field reports on private
AI-authored codebases drove eleven issues; this release closes all of them,
alongside the roadmap work that landed beside it.
Read this before upgrading. Defender discovery is now mock-aware: a test
that vi.mocks or jest.mocks the subject can detect no fault in it and is
no longer counted. Claims defended only by such tests therefore move to
NOCOVER, which is the truth those verdicts always should have told. Expect
new findings on the first probe after upgrading, and re-freeze your baseline
once you have read them — they are not regressions in your code, they are
blind spots that were previously invisible.
Added
-
testguard mcp— the operating loop, served over the Model Context
Protocol on stdio, so it survives a change of agent harness. The loop has
lived in a Claude Code skill and a session-start hook, and both vanish the
moment the harness is Cursor, Codex or whatever comes next. Five tools,
all read-only:testguard_status,testguard_brief,
testguard_claims,testguard_evidenceandtestguard_next_command.
No tool runs a probe — a probe is long-running, budgeted and the
person should see it happen, sonext_commandreturns the shell line
instead; and no tool writes a file, because editing a claims file through
a connector would defeat recording every fault edit. The server declares
only atoolscapability, and a tool that throws returns anisError
tool result rather than a protocol error, so one bad call never costs the
connection.
Hand-written against a pinned protocol version rather than built on the
official SDK: this tool keeps its single exact-pinned runtime dependency,
and the surface is three methods that will not grow. The tests speak the
wire format to a real child process and check every tool against the CLI's
own--jsonoutput, so the two cannot drift.
init --mcpprints the config for Claude Code, Cursor and Codex — printed,
never written, because a harness config is the person's file. Self-claims
TG-MCP-IS-READ-ONLYand
TG-MCP-TOOL-FAILURE-IS-NOT-A-PROTOCOL-ERROR. (#30) -
Contention detection and
--serial(#26).probelooks for other test
runners before the first run, warns naming their pids, and records them on
the evidence (run.contention) — a contended machine turns a slow suite
into aTIMEOUTorFLAKY-DEFENDERverdict about the load, not the claim,
and a later reader of that verdict can now tell.--serialruns one test
file at a time (--no-file-parallelism,--runInBand,--workers=1) and
is recorded asrun.serial. Best effort by design: an unreadable process
list is no detection and never fails a probe. -
claims --since <ref>— a claim that disappeared was invisible to every
other check:probeverifies what is there,gatesees the file covered by
another claim,statussays clean. Since deleting a claim is cheaper than
weakening its fault (whichchangedFaultsalready surfaces), removal is now
reported:removed-claimandremoved-faultgate, a rename that keeps the
statement verbatim isrenamed-claimand does not, and where evidence exists
the finding names the verdict the claim last had. Excused by aclaimentry
intestguard.ignore.jsonwith a reason, with expiry honoured as everywhere
else. This repository's CI runs it on its own pull requests. Found the hard
way: a--theirsconflict resolution dropped two self-claims during this
release and nothing noticed. Self-claimTG-REMOVED-CLAIM-IS-A-FINDING. (#57) -
testguard replay --since <range>— would this suite have caught the
bugs that already escaped? For each fix commit in the range (one that
changes source and a test together), it reverts only the source to the
parent in a scratch worktree, removes the test the fix shipped, and runs the
tests that import the reverted code:caught,blind,nocover,flaky
orunverifiable. One patch counts once (git patch-id). It reports and
never gates — a bug that escaped is history, not a regression in this
change. An injected fault is one somebody thought of; a bug that shipped is
ground truth, with no equivalent-mutant argument to have about it. -
Fault-class labelling and the first calibration. Each replayed bug is
labelled with the injected-fault class its diff most resembles — the
scaffold producers read in reverse, deterministic,otherrather than a
guess — and acalibrationdocument is written beside the replay one:
per class, the share of real escaped bugs the suite missed, with a Wilson
interval and n. Onlycaughtandblindcarry information; the rest are
excluded from both sides.calibration.schema.jsonhas had no producer
until now. The open question it exists to answer — does a calibration
learned on a repository with history transfer to a greenfield one — stays
open; this is the instrument, not the answer. (#35) -
New spec kind
replay(replay.schema.json), with semantic rules:
caughtneeds every run to fail by assertion,blindneeds every run to
pass,flakyneeds runs that disagree,nocovercannot have run tests,
and a duplicate patch-id is rejected. Conformance example plus two
must-reject documents. Self-claims
TG-REPLAY-FLAKY-IS-NEVER-CAUGHT,TG-REPLAY-DEDUPES-BY-PATCH,
TG-CALIBRATION-EXCLUDES-UNINFORMATIVEandTG-LABEL-NEVER-GUESSES. -
baseline --restamp(#25). A baseline frozen from--include-dirty
evidence now records the snapshot commit and, once you commit, a clean
probeplus--restampmoves itsheadto that commit — only when the
fingerprints are identical and the tree is clean; a frozen contract is
never silently rewritten.statusgains informationalnotes(never a
state): a baseline frozen from a snapshot or a dirty tree that predates
HEAD, or one whose head is not an ancestor of HEAD. -
Mock-aware defender discovery (#20). A test file that
vi.mocks /
jest.mocks the target cannot detect any fault in it and is no longer a
discovered defender;NOCOVERnow means "no test imports this source
without mocking it". The evidence lists such files under
defenders.mocking(a declared defender that mocks the subject stays,
and is listed, as a broken evidence chain);claimsprints the split
(18 import · 16 mock · 2 can detect). -
mocked-never-asserted(#21): a static signal for a test that mocks
the target and neverexpect(...)s anything imported from it — the exact
signature of an escaped bug in a field report. Recorded on the evidence
(defenders.signals), printed byclaims, carried into the brief's hints.
// unasserted: <why>above the mock silences it visibly
(unasserted-annotatedwith the reason). -
Alias resolution follows tsconfig
references(the Vite layout, where
pathslive intsconfig.app.json) and reads vite/vitestresolve.alias
as text, so importers from nested__tests__/directories are found. -
Both known-answer fixtures gain a claim whose only importing test mocks
the module: expectedNOCOVER, with the mocking file and the signal
recorded. -
testguard admit <test-file> --claim <ID>— the two-gate rule as a
named verb. Sugar overprobe --claim <ID> --include-dirty --no-escalate:
the named test must be a declared or discovered defender of the claim
(exit3otherwise, with thedefendedByline to add);ADMITTED
(exit0) only when every fault of the claim iskilledN/N on defenders
that were green N/N unmodified; anything else isNOT ADMITTED(exit1)
and names the first blocking fault with the hint the brief would give.
--fault <FID>judges one fault,--confirm 1gives a provisional
ADMITTED?,--jsonreturns{admitted, provisional, faults[], evidence, command}. Evidence goes to.testguard/evidence-partial.json; nothing new
in the evidence schema. The market's acceptance signal for a generated
test is "compiles, passes, raises coverage"; this one is "fails when the
claim is false", and it is now one command.status.nextforunproven
and the installed skill's fix loop point at it. Self-claim
TG-ADMIT-NEEDS-ALL-KILLED.
Changed
-
initinstalls the agent layer at the git root (#23): the skill,
the session-start hook and theAGENTS.mdsection go where agent
sessions run; the.gitignorelines stay beside the claims file. A
second project in the same repository adds a hook line and an
AGENTS.mdbullet;--herekeeps the old placement. A written file
that.gitignoreswallows is reported (exit 1), never offered for
commit. -
The session-start hook resolves an installed
testguardand contains no
form ofnpx(#24): the project'snode_modules/.bin, then the git
root's, thencommand -v testguard, then nothing — it exits 0 with no
output rather than break a session. Any earlier npx hook,-yor
--no-install, is replaced on the nextinit. The brief's first line
names the install that answered.--no-installis quiet, not offline, which is why it is gone too:
measured,npx --no-install --loglevel=http testguard-cli --versionin a
project with nothing installed logs
npm http fetch GET 200 https://registry.npmjs.org/testguard-cli, and
against an unreachable registry it exits non-zero. npm resolves the
packument before deciding not to install. -
probehonours an explicit--ref(even--ref HEAD) when defender or
target files are dirty, and gains--ignore-dirtyfor the implicit HEAD
(#19). Both warn with the file names and record them in the evidence as
repo.ignoredDirty(spec: new optional field; a snapshot run can never
carry it). The refusal stays for the implicit HEAD without the flag — that
is the silent-mismatch trap the check exists for. -
CI: TestGuard's self-probe runs on one Node leg instead of three, without
escalation, and restores the previous run's evidence from the cache so
unchanged claims reuse their verdicts. The verdicts do not depend on the
Node minor; the from-scratch probe was ~14 minutes per leg.
Fixed
-
The README described a hook the code no longer writes. Three passages
still documented thenpx --no-installfallback that this release removes,
one of them repeating the "never fetches from the network" wording the
measurement above disproves — in the security-relevant paragraph, in a
release about detection power. Corrected, and now checked mechanically:
every||-chainedbrief --textcommand in the README must be one
hookCommand()actually emits, and the phrase "falls back to npx" is
banned outright. Both regressions were verified to fail the check before it
was accepted. Self-claimTG-README-HOOK-MATCHES-THE-CODE. -
A claim can be defended and still over-promise, and no probe can catch
that.TG-INIT-HOOK-NO-NETWORKwas killed 3/3 on every run while the
code it guarded could still reach the registry: the fault matched the
code, the code matched the test, and the statement was the thing that
was wrong. Fault injection measures whether a test would notice the code
changing; it cannot measure whether the sentence a human wrote is true of
the world. That is the standing limit of this method, and the only remedy
is reading claims against reality — which is what happened here. The
statement and its fault were corrected together. -
The session-start hook now contains no form of
npx. It fell back to
npx --no-install, and--no-installis quiet rather than offline:
measured,npx --no-install --loglevel=http testguard-cli --versionin a
project with nothing installed logs
npm http fetch GET 200 https://registry.npmjs.org/testguard-cli, and
against an unreachable registry it exits non-zero. npm resolves the
packument before deciding not to install. The fallback is now a
command -vlookup, which covers a global install with no network at all.
TG-INIT-HOOK-NO-NETWORKwas defended while its statement over-promised;
the statement and its fault are corrected together. -
The PATH branch is braced.
a || b && cbinds as(a || b) && cin sh, so
the unbraced form printed the brief twice whenever the local binary
succeeded. The tests now execute the hook in all three cases (local, PATH,
neither) rather than matching its text. -
initrecognises an earliernpx --no-installhook as well asnpx -y,
and its which-project check no longer mistakesbrief --text 2>/dev/null
for a hook belonging to a directory — that misfire meant a legacy root hook
was never replaced. -
replayscopes to the project directory. In a monorepo a fix commit
routinely touches several packages; the harness kept the out-of-project
files in the revert, so every cross-package fix came back
revert-did-not-apply— a tooling failure wearing the costume of a
verdict. Found on a real corpus, where it was two of the first three
commits. A commit with no source-and-test pair inside the project is no
longer a candidate at all. Self-claimTG-REPLAY-SCOPES-TO-THE-PROJECT. -
replayreverts a file the fix ADDED by removing it, and calls a commit
whose source is entirely newno-prior-versionrather than a failed revert.
A fix routinely adds a helper as well as changing a module, and a file that
did not exist at the parent cannot be checked out of it: on a real corpus
this turned nineteen of forty commits intorevert-did-not-apply— a
tooling failure reading as a verdict and hiding every real result behind it.
An addition is not a bug the suite could have caught, so it never enters a
calibration. Self-claim
TG-REPLAY-ADDED-FILE-IS-REMOVED-NOT-CHECKED-OUT. -
replay --out <path>now writes the calibration beside it rather than into
the project's.testguard/: the two documents are one result and splitting
them loses the pairing.