Skip to content

TestGuard v0.6.0

Choose a tag to compare

@github-actions github-actions released this 18 Sep 01:16
· 107 commits to main since this release
0bf0f64

Everything the field asked for. Two independent field reports on private
AI-authored codebases drove eleven issues; this release closes all of them,
alongside the roadmap work that landed beside it.

Read this before upgrading. Defender discovery is now mock-aware: a test
that vi.mocks or jest.mocks the subject can detect no fault in it and is
no longer counted. Claims defended only by such tests therefore move to
NOCOVER, which is the truth those verdicts always should have told. Expect
new findings on the first probe after upgrading, and re-freeze your baseline
once you have read them — they are not regressions in your code, they are
blind spots that were previously invisible.

Added

  • testguard mcp — the operating loop, served over the Model Context
    Protocol on stdio, so it survives a change of agent harness. The loop has
    lived in a Claude Code skill and a session-start hook, and both vanish the
    moment the harness is Cursor, Codex or whatever comes next. Five tools,
    all read-only: testguard_status, testguard_brief,
    testguard_claims, testguard_evidence and testguard_next_command.
    No tool runs a probe — a probe is long-running, budgeted and the
    person should see it happen, so next_command returns the shell line
    instead; and no tool writes a file, because editing a claims file through
    a connector would defeat recording every fault edit. The server declares
    only a tools capability, and a tool that throws returns an isError
    tool result rather than a protocol error, so one bad call never costs the
    connection.
    Hand-written against a pinned protocol version rather than built on the
    official SDK: this tool keeps its single exact-pinned runtime dependency,
    and the surface is three methods that will not grow. The tests speak the
    wire format to a real child process and check every tool against the CLI's
    own --json output, so the two cannot drift.
    init --mcp prints the config for Claude Code, Cursor and Codex — printed,
    never written, because a harness config is the person's file. Self-claims
    TG-MCP-IS-READ-ONLY and
    TG-MCP-TOOL-FAILURE-IS-NOT-A-PROTOCOL-ERROR. (#30)

  • Contention detection and --serial (#26). probe looks for other test
    runners before the first run, warns naming their pids, and records them on
    the evidence (run.contention) — a contended machine turns a slow suite
    into a TIMEOUT or FLAKY-DEFENDER verdict about the load, not the claim,
    and a later reader of that verdict can now tell. --serial runs one test
    file at a time (--no-file-parallelism, --runInBand, --workers=1) and
    is recorded as run.serial. Best effort by design: an unreadable process
    list is no detection and never fails a probe.

  • claims --since <ref> — a claim that disappeared was invisible to every
    other check: probe verifies what is there, gate sees the file covered by
    another claim, status says clean. Since deleting a claim is cheaper than
    weakening its fault (which changedFaults already surfaces), removal is now
    reported: removed-claim and removed-fault gate, a rename that keeps the
    statement verbatim is renamed-claim and does not, and where evidence exists
    the finding names the verdict the claim last had. Excused by a claim entry
    in testguard.ignore.json with a reason, with expiry honoured as everywhere
    else. This repository's CI runs it on its own pull requests. Found the hard
    way: a --theirs conflict resolution dropped two self-claims during this
    release and nothing noticed. Self-claim TG-REMOVED-CLAIM-IS-A-FINDING. (#57)

  • testguard replay --since <range> — would this suite have caught the
    bugs that already escaped? For each fix commit in the range (one that
    changes source and a test together), it reverts only the source to the
    parent in a scratch worktree, removes the test the fix shipped, and runs the
    tests that import the reverted code: caught, blind, nocover, flaky
    or unverifiable. One patch counts once (git patch-id). It reports and
    never gates — a bug that escaped is history, not a regression in this
    change. An injected fault is one somebody thought of; a bug that shipped is
    ground truth, with no equivalent-mutant argument to have about it.

  • Fault-class labelling and the first calibration. Each replayed bug is
    labelled with the injected-fault class its diff most resembles — the
    scaffold producers read in reverse, deterministic, other rather than a
    guess — and a calibration document is written beside the replay one:
    per class, the share of real escaped bugs the suite missed, with a Wilson
    interval and n. Only caught and blind carry information; the rest are
    excluded from both sides. calibration.schema.json has had no producer
    until now. The open question it exists to answer — does a calibration
    learned on a repository with history transfer to a greenfield one — stays
    open; this is the instrument, not the answer. (#35)

  • New spec kind replay (replay.schema.json), with semantic rules:
    caught needs every run to fail by assertion, blind needs every run to
    pass, flaky needs runs that disagree, nocover cannot have run tests,
    and a duplicate patch-id is rejected. Conformance example plus two
    must-reject documents. Self-claims
    TG-REPLAY-FLAKY-IS-NEVER-CAUGHT, TG-REPLAY-DEDUPES-BY-PATCH,
    TG-CALIBRATION-EXCLUDES-UNINFORMATIVE and TG-LABEL-NEVER-GUESSES.

  • baseline --restamp (#25). A baseline frozen from --include-dirty
    evidence now records the snapshot commit and, once you commit, a clean
    probe plus --restamp moves its head to that commit — only when the
    fingerprints are identical and the tree is clean; a frozen contract is
    never silently rewritten. status gains informational notes (never a
    state): a baseline frozen from a snapshot or a dirty tree that predates
    HEAD, or one whose head is not an ancestor of HEAD.

  • Mock-aware defender discovery (#20). A test file that vi.mocks /
    jest.mocks the target cannot detect any fault in it and is no longer a
    discovered defender; NOCOVER now means "no test imports this source
    without mocking it". The evidence lists such files under
    defenders.mocking (a declared defender that mocks the subject stays,
    and is listed, as a broken evidence chain); claims prints the split
    (18 import · 16 mock · 2 can detect).

  • mocked-never-asserted (#21): a static signal for a test that mocks
    the target and never expect(...)s anything imported from it — the exact
    signature of an escaped bug in a field report. Recorded on the evidence
    (defenders.signals), printed by claims, carried into the brief's hints.
    // unasserted: <why> above the mock silences it visibly
    (unasserted-annotated with the reason).

  • Alias resolution follows tsconfig references (the Vite layout, where
    paths live in tsconfig.app.json) and reads vite/vitest resolve.alias
    as text, so importers from nested __tests__/ directories are found.

  • Both known-answer fixtures gain a claim whose only importing test mocks
    the module: expected NOCOVER, with the mocking file and the signal
    recorded.

  • testguard admit <test-file> --claim <ID> — the two-gate rule as a
    named verb. Sugar over probe --claim <ID> --include-dirty --no-escalate:
    the named test must be a declared or discovered defender of the claim
    (exit 3 otherwise, with the defendedBy line to add); ADMITTED
    (exit 0) only when every fault of the claim is killed N/N on defenders
    that were green N/N unmodified; anything else is NOT ADMITTED (exit 1)
    and names the first blocking fault with the hint the brief would give.
    --fault <FID> judges one fault, --confirm 1 gives a provisional
    ADMITTED?, --json returns {admitted, provisional, faults[], evidence, command}. Evidence goes to .testguard/evidence-partial.json; nothing new
    in the evidence schema. The market's acceptance signal for a generated
    test is "compiles, passes, raises coverage"; this one is "fails when the
    claim is false", and it is now one command. status.next for unproven
    and the installed skill's fix loop point at it. Self-claim
    TG-ADMIT-NEEDS-ALL-KILLED.

Changed

  • init installs the agent layer at the git root (#23): the skill,
    the session-start hook and the AGENTS.md section go where agent
    sessions run; the .gitignore lines stay beside the claims file. A
    second project in the same repository adds a hook line and an
    AGENTS.md bullet; --here keeps the old placement. A written file
    that .gitignore swallows is reported (exit 1), never offered for
    commit.

  • The session-start hook resolves an installed testguard and contains no
    form of npx
    (#24): the project's node_modules/.bin, then the git
    root's, then command -v testguard, then nothing — it exits 0 with no
    output rather than break a session. Any earlier npx hook, -y or
    --no-install, is replaced on the next init. The brief's first line
    names the install that answered.

    --no-install is quiet, not offline, which is why it is gone too:
    measured, npx --no-install --loglevel=http testguard-cli --version in a
    project with nothing installed logs
    npm http fetch GET 200 https://registry.npmjs.org/testguard-cli, and
    against an unreachable registry it exits non-zero. npm resolves the
    packument before deciding not to install.

  • probe honours an explicit --ref (even --ref HEAD) when defender or
    target files are dirty, and gains --ignore-dirty for the implicit HEAD
    (#19). Both warn with the file names and record them in the evidence as
    repo.ignoredDirty (spec: new optional field; a snapshot run can never
    carry it). The refusal stays for the implicit HEAD without the flag — that
    is the silent-mismatch trap the check exists for.

  • CI: TestGuard's self-probe runs on one Node leg instead of three, without
    escalation, and restores the previous run's evidence from the cache so
    unchanged claims reuse their verdicts. The verdicts do not depend on the
    Node minor; the from-scratch probe was ~14 minutes per leg.

Fixed

  • The README described a hook the code no longer writes. Three passages
    still documented the npx --no-install fallback that this release removes,
    one of them repeating the "never fetches from the network" wording the
    measurement above disproves — in the security-relevant paragraph, in a
    release about detection power. Corrected, and now checked mechanically:
    every ||-chained brief --text command in the README must be one
    hookCommand() actually emits, and the phrase "falls back to npx" is
    banned outright. Both regressions were verified to fail the check before it
    was accepted. Self-claim TG-README-HOOK-MATCHES-THE-CODE.

  • A claim can be defended and still over-promise, and no probe can catch
    that.
    TG-INIT-HOOK-NO-NETWORK was killed 3/3 on every run while the
    code it guarded could still reach the registry: the fault matched the
    code, the code matched the test, and the statement was the thing that
    was wrong. Fault injection measures whether a test would notice the code
    changing; it cannot measure whether the sentence a human wrote is true of
    the world. That is the standing limit of this method, and the only remedy
    is reading claims against reality — which is what happened here. The
    statement and its fault were corrected together.

  • The session-start hook now contains no form of npx. It fell back to
    npx --no-install, and --no-install is quiet rather than offline:
    measured, npx --no-install --loglevel=http testguard-cli --version in a
    project with nothing installed logs
    npm http fetch GET 200 https://registry.npmjs.org/testguard-cli, and
    against an unreachable registry it exits non-zero. npm resolves the
    packument before deciding not to install. The fallback is now a
    command -v lookup, which covers a global install with no network at all.
    TG-INIT-HOOK-NO-NETWORK was defended while its statement over-promised;
    the statement and its fault are corrected together.

  • The PATH branch is braced. a || b && c binds as (a || b) && c in sh, so
    the unbraced form printed the brief twice whenever the local binary
    succeeded. The tests now execute the hook in all three cases (local, PATH,
    neither) rather than matching its text.

  • init recognises an earlier npx --no-install hook as well as npx -y,
    and its which-project check no longer mistakes brief --text 2>/dev/null
    for a hook belonging to a directory — that misfire meant a legacy root hook
    was never replaced.

  • replay scopes to the project directory. In a monorepo a fix commit
    routinely touches several packages; the harness kept the out-of-project
    files in the revert, so every cross-package fix came back
    revert-did-not-apply — a tooling failure wearing the costume of a
    verdict. Found on a real corpus, where it was two of the first three
    commits. A commit with no source-and-test pair inside the project is no
    longer a candidate at all. Self-claim TG-REPLAY-SCOPES-TO-THE-PROJECT.

  • replay reverts a file the fix ADDED by removing it, and calls a commit
    whose source is entirely new no-prior-version rather than a failed revert.
    A fix routinely adds a helper as well as changing a module, and a file that
    did not exist at the parent cannot be checked out of it: on a real corpus
    this turned nineteen of forty commits into revert-did-not-apply — a
    tooling failure reading as a verdict and hiding every real result behind it.
    An addition is not a bug the suite could have caught, so it never enters a
    calibration. Self-claim
    TG-REPLAY-ADDED-FILE-IS-REMOVED-NOT-CHECKED-OUT.

  • replay --out <path> now writes the calibration beside it rather than into
    the project's .testguard/: the two documents are one result and splitting
    them loses the pairing.