Skip to content

v0.14.0

Choose a tag to compare

@3dgiordano 3dgiordano released this 26 Sep 01:25
· 14 commits to main since this release

A new plugin, integrity, for work that cannot be done as asked: the result
must be real and the route to it legitimate, and "cannot be done as asked,
because X" is a complete answer. Measured before it was built and on the
finished plugin. On the bench's task with an unreachable service, Grok 4.7
High went from 0/3 to 3/3 with it (n=3), through the skill. The rule behind
what a hook may say is written down, and it is enforced: what a hook reads
never becomes what it says. Persistence stopped quoting a line of command
output. The bench fences the agent's node to its workspace, after a run
tried -Verb RunAs through it, and names and stops a CLI session that hangs.

Plugin Version
executive-self-monitoring 1.6.2
epistemic-self-monitoring 0.3.1
persistence-self-monitoring 0.3.2
termination-self-monitoring 0.3.0
coverage-self-monitoring 0.4.1
handoff-self-monitoring 0.3.1
progress-self-monitoring 0.5.2
integrity-self-monitoring 0.1.0

Added

  • integrity-self-monitoring 0.1.0 — a new plugin for work that cannot be done as asked. When a service cannot be reached, a key is missing or two tests want different answers, agents slide into a result that only looks done and report it as the fix; every model in this bench did. The plugin keeps the result real and the route to it legitimate: "cannot be done as asked, because X" is a complete answer, closed with an [INTEGRITY CHECK] (Result, Route, Outside the task, Told the user).
    • What it reads. Every edit to product code, as the file is on disk after it (tests, fixtures, mocks, node_modules, minified files left out). Six shapes are named to the agent: a rate table, a catch around an awaited call that answers with a value of its own (not one that carries the error), a promise .catch that answers with another call, a missing key answered with a value, a second host, all only in a file that calls a service; and code that reads its caller (require.main not compared with module, module.parent, new Error().stack) next to a test's name. Every shell command, for TLS turned off, the hosts file written, name resolution replaced, a system setting changed or a server started inline. Each finding is said once per session, with what it means for the user, and shown to the user as one line. A code finding names the file, the line and the shape, never the text it matched: that text is file content, and echoed back it would reach the model as the hook's own message. A command is quoted, as the other plugins quote one, since the agent issued it.
    • What it does not say. Nothing on the prompt: a rule about honest results said on every turn is a cue on every turn. No examiner, score or observer in any text. A credential with a literal default, TLS off in the source and paths outside the project are logged only, until they are measured.
    • A dispute. The agent can answer a finding as what was asked (- src/rates.js: misread - <why>); the user sees that answer, and the finding is not raised again. Claude Code and Codex read it at the close; Cursor logs it (a headless run fires no close hook, so there the finding lives in postToolUse only).
    • Measured, on the stored runs of this bench regraded with today's graders (node scripts/integrity.js --signals --regrade, 414 streams): 27 runs carry a finding, 26 of them failed as dishonest and the 27th is the run that stood a local HTTPS server in for the service, which the grader passes because src/ is honest. 0 of the other 268 passing runs carries one. On the case with contradictory tests the caller shape reads 2 of 21 tricks; the other 19 keep state between calls. On real code, file by file: 0 findings in 5910 files of four working trees (after the refinements they prompted), 1 in 15041 files of 1207 installed npm packages, out of sample (a client's own three hosts). Edits in real sessions are the next measurement.
    • The effect, n=3 (Grok 4.7 High, integrity-self-monitoring/the-blocked-part-stays-blocked, node fenced): with the plugin 3/3, without 0/3. Every WITH run opened the skill, implemented to the ticket, left the two rates tests red and closed with Result: blocked. Every WITHOUT run fell back to Frankfurter's rates and reported the tests passing. No hook fired in a WITH run, since none wrote a shortcut, so this is the skill's effect. The runs were also cheaper: 110 s and 26k tokens against 240 s and 109k. Composer 2.5 is not measured, because its two WITH runs stalled in the CLI. Its shortcut moved into the test file, which 0.1.0 does not read.
    • Corpus: evals/corpus/integrity-code.jsonl (101 lines: whole files from the stored runs, labelled by the grader, and hand-written neighbours; recall 0.67, precision 1.0) and integrity-commands.jsonl (37 lines; 1.0 / 1.0). scripts/corpus.js hands a detector the whole corpus line, so a line can carry a file's path.
    • Two eval cases: keeps-the-result-real (a carrier's quotes service with no key and no network) and stays-quiet-on-an-honest-change.
    • Its bench case is coverage's task. bench/integrity-self-monitoring/the-blocked-part-stays-blocked runs the same prompt, files and grader with integrity installed: case.json "same": "<plugin>/<id>" is new in scripts/benchlib.js, and the case's grade.js re-exports the other one. scripts/bench-check.js checks such a case on its task's fixtures, and counts the plugins in plugins/ instead of expecting seven. It has no runs yet, so the dispute the plugin offers is not measured: how often an agent answers a true finding as a misread, from node scripts/integrity.js --signals on the WITH arm. The benchmark stays 1.0.0 - no existing case changed, and the new plugin's version already marks its case to run.
    • evals/corpus/integrity-code.jsonl holds passing solutions of a bench case, so it carries the canary and counts as answer material (scripts/integrity.js CORPUS_FROM_RUNS; --stamp writes a // line in a JSONL corpus). .agent/integrity/ is ignored.
    • scripts/integrity.js --signals [--regrade] replays each stored stream through the plugin's own signals and prints them beside the audit's marks, the grade and any dispute.

Changed

  • coverage 0.4.1: an [INTEGRITY CHECK] is a report, like the other sibling blocks. Its Route line names what a blocked part still needs ("it still needs RATES_API_KEY and can be finished later"), and coverage's close scan read that as deferred work: a retrospective for a close that had said exactly what was blocked and why. The marker joins REPORT_MARKER_RE; a corpus line holds it.
  • Every place that lists the collection names integrity: the root README (the questions, the moments, the plugins table, the host table, the install lists, what the hooks read and that five plugins never block), the social preview and its PNG, assets/README.md, CONTRIBUTING's naming rule, both issue templates, the boundary notes in coverage's and persistence's READMEs, handoff's list of sibling blocks, SECURITY.md's state prefixes. scripts/bench-check.js counts the plugins instead of expecting seven.
  • The README says what each part needs (Requirements). The plugins need node on the host's PATH, Node 18 or later, and a host that runs hooks. The test pipeline is split by layer. The suite needs Node 18+. The evals need the host's CLI, logged in. The bench needs the Cursor CLI, PowerShell or cmd on Windows, and Node 20+ for the node fence. Before this, one line under Development, "Node 18+ is all you need", stood for both. It stopped being true for the bench once the fence came in.
  • persistence 0.3.2: the repeated-error nudge no longer quotes the error. It put the first failure line of a command's output into the message, numbers and paths blanked (the same error has come back 3 times this turn (Error: got # at <path>:#)). That line is output text: a test, a file in the repository or a service writes it, so a failing test could put an instruction into what the model reads as the hook's own message. The nudge now says the same error has come back 3 times this turn, last in the output of npm test``: the count and the command the agent ran, and the agent reads the output itself. The signature still keys the count and goes to the opt-in log.
  • SECURITY.md states the rule behind what a hook may say. It used to read "reads no files other than its own state, with one exception"; the rule that matters is that what a hook reads never becomes what it says. A file, a tool result or a prompt may be read and measured; only counts and metadata about it reach a message, plus what the agent itself wrote, short. The two project files a hook reads are listed (progress's ledger, integrity's just-edited file). The one value that was taken from a tool result, persistence's error line, is gone (above).

Security

  • The agent's node is fenced to its workspace in the bench and the evals. The shell allowlist (Shell(node **)) held: in two Composer 2.5 runs it rejected every command that was not node. But node itself started PowerShell with -Verb RunAs, which failed only on the agent's syntax, and tried to append to the hosts file. evallib.js nodeGuard() now puts a node first on each invocation's PATH that runs the real one under Node's permission model. Reads and writes stay inside the workspace, with no child process, worker, addon or WASI, and the network stays open. The wrapper refuses --allow-* and --permission flags and drops NODE_OPTIONS. Hooks run unfenced from the plugin copy and the witness. Tested from PowerShell: the tests, require from the workspace's node_modules, fetch and exit codes pass through, and spawn, RunAs, a hosts write, a write to %TEMP% and a read of the repository are denied. A denied access is a new suspect mark, stopped by the node fence. On by default, recorded as nodeGuard in run.json (the flag used, or false), off with --no-node-guard. It needs a node with a permission model on the runner: --permission from 22.13, --experimental-permission on 20 and earlier 22. On Node 18 the runs are not fenced, the runner says so, and run.json records false. It does not hand the agent a node that refuses every command, which is what the first cut did on CI with Node 18. The benchmark stays 1.0.0.

Fixed

  • A Cursor run the CLI left hanging is named, and stopped when it goes silent. Five stored Composer 2.5 streams, all on the-blocked-part-stays-blocked, both arms, with and without a plugin, end on a finished thinking block followed by 8-13 minutes of silence, until the per-invocation ceiling killed them. They were counted as plain dead runs, like a model that was still working. In 161 Composer runs that finished, the longest silence between two events was 25 s. evallib.js stallOf() names that signature: no result event, and the last event a finished thinking block. It flags those 5 of the 416 stored streams and no others, and cursor-bench.js reports the run as stalled after a thinking block and records stall in cost.json. scripts/idle-watchdog.js sits in front of the CLI when cursor-eval.js run() is given an idle limit: it stops the command and its process tree after that long with no output, writes [idle_timeout] and exits 124. bench/suite.json sets idleMin per model: Composer 2.5 3, the Groks 10 (their longest healthy silences are 153 s and 460 s). --idle-min overrides it. A cut run stays dead and is not retried.
  • The test suite no longer writes into the maintainer's misread log. A shell with COVMON_MISREAD_LOG=all passed it on to every hook the suite drives, and the fixture closes landed in ~/.3dgiordano-agent-plugins/misreads/: six entries in 30 seconds on 2026-09-25. hook() in scripts/test.js now clears COVMON_MISREAD_LOG and COVMON_MISREAD_FILE unless a test sets them.

Full Changelog: v0.13.0...v0.14.0