Skip to content

Choose a tag to compare

@mavericksea-ai mavericksea-ai released this 06 Oct 04:02

What's new

  • Start from a skill you already have. In Claude Code, /driftproof:init or /driftproof:run on a
    skill that already exists now offers the first-run path: Claude drafts test cases from the skill with
    you, adds them to the skill's folder only after you say yes, runs a quick first look, and opens a
    results page from disk. /driftproof:start is the same path under its own name: with no folder it
    lists the skills it finds and whether each has test cases. The README's Quickstart now leads with
    this path, and the scaffold for a new skill is the second path.
  • A quick first run. driftproof run <skill> --quick runs a short smoke test: 2 judge samples, 4
    calls at a time and at most 5 test cases. Its receipt says it cannot produce a verdict, and its output
    prints no lift and no result word. Use it to see that the setup works; run without --quick for a
    result you can rely on.
  • Add test cases to a skill that exists. /driftproof:init can add one new file, evals/evals.json,
    to a skill folder, only after you say yes. It never changes your SKILL.md or any other file in the
    folder.
  • Run a skill you wrote outside a git repository. /driftproof:run takes --trust-outside-repo,
    which you give once you confirm you wrote the skill. A missing isolation account now gets one plain
    message that names this route.
  • Fewer real answers are thrown out as lost. A draw is now counted as lost only when its reply
    points at a file the judge was not shown. File lists in commit or pull request text, answers inside a
    fenced block and first-person advice are no longer read as lost answers. In --capture files mode, a
    model that writes its answer to a file and replies with nothing is graded on the file, and in text
    mode the tools that can write files or run code are off.
  • Clearer words when the cases disagree. When the test cases disagree so much that more draws
    would not settle the result, the result now says "More cases, not more draws, are needed to conclude
    at this effect floor". "Not enough draws" is kept for the case where more draws would help.
  • Safer pull request comments and job summaries. The Action's comment and summary, and the stale
    Action's summary and issue, now escape HTML, Markdown links and URLs in every string that comes from a
    receipt or a folder, so a case id or file name that holds them reads as the characters it is. Mentions
    and issue references are not covered yet: see Known issues.
  • The stale check reads the capture mode. driftproof stale treats the capture mode (text or
    files) as one more thing that can make a receipt out of date, beside the model, the skill and the
    harness.
  • Better docs. The README now explains how to read what a run prints, what --capture changes and
    what /driftproof:run takes. driftproof view labels a receipt from before the answered_by field
    existed as "Recorded before answered_by existed", where it said "Not measured".
  • Dependency and code scanning fixes. The lockfile's fast-uri is 3.1.8, and the CodeQL alerts
    open at the cut are fixed in code.
  • Published reports gain dated notes. Reports 001 to 005, 007 and 008 each gain a dated note, and
    Report 007 gains three dated corrections to its judge cost figures. No published number is changed.

What may change for you

  • /driftproof:init on a skill that already exists, with no flags, now refuses and writes
    nothing.
    It exits with 2 and says it would add evals/evals.json to the skill, and that it adds a
    file to a skill folder only with your yes. To add drafted test cases to a skill that exists, give the
    drafted cases with --cases <file> and your yes with --confirm-write:
    /driftproof:init <skill folder> --cases <file> --confirm-write. --confirm-write alone adds
    placeholder example cases, and --cases without it is refused. A skill that already has test cases
    gets its quick run through /driftproof:start <skill folder>, which runs them and adds nothing.
    The command line's driftproof init <folder> on a folder that holds a SKILL.md adds
    evals/evals.json only, never your SKILL.md, and no longer writes a .driftproofrc there.
  • The stale check now says "advisory" for a receipt with no capture field. Every receipt made
    before receipt spec 0.10 has none. A scheduled stale Action on such a receipt now reads advisory
    (exit 3) where it read current, and with strict it reads stale and opens an issue. To clear it, run
    again: a receipt from this release records its capture mode.
  • The README's Quickstart link has a new anchor. The heading is now "Quickstart: receipt for your
    own skill", so a link that ends #quickstart--receipt-for-your-own-skill-in-10-minutes needs
    #quickstart-receipt-for-your-own-skill. The published 0.13.0 page keeps the old anchor.
  • Some words on the view page, the pull request comment and the job summary change. For a receipt
    from before answered_by existed, the label moves from "Not measured" to "Recorded before
    answered_by existed". 129 of the 262 receipts in this repository are in that state, and no verdict
    moves. The site's receipt pages still say "Not measured" for them. Three receipt pages that show a
    one-case verdict now say it is for one case, on one test task.
  • Receipts written by this release are receipt spec v0.11. v0.11 adds an optional run.preset,
    "quick" on a smoke run's receipt, and such a receipt cannot read as a measured result. v0.10 is
    frozen as spec/receipt.v0.10.schema.json, and every earlier receipt still validates against its
    own version.
  • No published receipt, number or verdict changes. Reports 001 to 005, 007 and 008 gain a dated note,
    Report 009's opening gains one line, and Report 007's three judge cost figures are left as published,
    with a dated correction under each. No receipt file changes.

Known issues

  • Cases from a skill that does not exist. On the command line, driftproof init <new folder> --cases <file> --drafted-from-skill creates a stub SKILL.md and labels the suite and each case as
    drafted from the skill and approved by the user, though there was no skill to draft from. The
    plugin never takes this route.
  • A lost-answer reading still has edges. A bare "Created report.md." beside a pleasantry, a
    reply that opens with a title, and some code names that end in a file extension can still be read the
    wrong way. They are named in the engineering log in RELEASES.md.
  • Mentions and issue references still fire. Measured on GitHub on 6 Oct 2026: in a comment, an
    escaped @name still renders as a user mention, and an escaped #1 or GH-1 still renders as an issue
    link. HTML, Markdown links, bare URLs and character references render as plain text. The fix is planned
    for the next release.
  • Backslashes in the stale summary. A path or model id that holds a backslash shows it doubled in
    the stale summary's code cells.
  • driftproof view and a bad capture value. A .driftproofrc in the working directory with
    "capture": "both" makes view list the receipt as not read, where stale and run exit 2.
  • One check clears at the publish. The plugin's check that its package matches the published npm
    package can read only after 0.14.0 is on npm.

Upgrade

  • npm: npx driftproof@0.14.0. The Action: driftproofhq/driftproof@v0.14.0, and the stale Action
    driftproofhq/driftproof/stale@v0.14.0.
  • Claude Code plugin: claude plugin update driftproof@driftproofhq. The plugin runs npx driftproof@0.14.0, and /driftproof:start needs a runner at that version or later.
  • Receipts from earlier versions still validate. A workflow that does not set the new options behaves
    as on 0.13.0, except for the changes listed under What may change for you.