Repository navigation
v0.3.0
agentseam 0.3.0
The evidence release. Claims on the capability matrix now carry their own evidence records, grades are capped by the kind of evidence behind them, a witnessed run can be frozen and replayed in CI, and Claude Code's three blocking events are witnessed on one real build. Eight adapter defects found by the vendor-truth review are fixed. No wire-format break: every consumer's hook config installs the same.
Upgrading from 0.2.1: pip install agentseam==0.3.0. Consumers that vendor bundles (bundler.bundle(agent)) will see them regenerate — the adapter fixes below are in the bundle, so re-run your bundling step and commit the result.
Highlights
- Claude Code's
prompt_submitandstopcells are witnessed at 2.1.263, the waypre_toolalready was. Both gates re-run against the real CLI, twice each, identical results. Two facts the run established: a headless{"decision": "block"}at Stop refuses the agent permission to finish and sends it round again, and Claude Code caps that loop at eight re-fires before ending the turn regardless. - Recorded driver — freeze a witnessed run, replay it in CI. One immutable file per witnessed
(agent, version)underdata/recordings/, read by the package viaagentseam.recordings.tools/experiment.py run --recordfreezes a real-agent run;--driver recordedreplays it through the same classifier a live run uses, with no process launched — the sevenclaude_code@2.1.263trials reproduce in under a second. The basis chain is now claim → recording → live run. - Per-claim evidence, and grading capped by basis. Every asserted cell field (
block,rewrite,fail_mode) carries its own{basis, date, version?, test | method}record.matrix.enforcement_level()no longer returns a grade its basis cannot support — avendor-docscell asserting fail-closed gradesbest-effort, neverenforced(the defect that shipped as chock#89). Ceiling table:matrix_terms.GRADE_CEILING. No current row's grade changes; the gap is closed defensively. - Three new optional cell fields for behaviours the experiment kit already measures:
silence_means,timeout_fail_mode,unknown_verb_means. Absence is not a claim. A recognised field a cell does not carry now readsunasserted, neverunrecordedby omission. - An
escalatetrial, answered in each agent's own dialect (Cursor spells itask). New measured fieldescalate_means, which can readprompted— the run ended waiting on an answer nobody gave. - The grade cap honours
verified.observed. Alive-run-partialrow can no longer let an unobserved event back a grade as high asenforced; unobserved events fall back to the row'sverified.fallback_basis(defaultvendor-docs). 11 of 91 claimed pairs change basis; none changes grade.
Adapter fixes (vendor-truth review)
Frozen wire output moves by a handful of bytes in total; each is called out in the changelog.
- Cursor — the prompt gate no longer surfaces an allow's own rationale in the UI on every submitted prompt, and explains a refusal that was really a degraded rewrite. A
Decision.rewrite(None, …)atpreToolUseis no longer blamed on the one gate that can express a rewrite.respond()infers an unnamededits[]payload's event the same wayparse()does, so the two can no longer diverge into a permission verdict at a read-only event. - Kimi Code — a payload that self-identifies as Kimi but names an event Kimi has not mapped now reaches the caller as
UNKNOWNinstead of "unrecognized payload" (claims.accept_any_name). A KimiPermissionRequestis no longer also claimed by Devin (claims.reject_client_types), which had left it unidentified. A control character in an installed command no longer renders aconfig.tomlthe vendor refuses to load at all — the full TOML basic-string escape set is emitted and pinned by atomllibround-trip. - Devin — the degraded-rewrite note at
UserPromptSubmit/Stopno longer describes a tool call those events don't have. - Grok and Kimi Code —
PostCompactno longer masquerades as canonicalpre_compact; a handler that snapshots context before compaction no longer also fires after it.
Instrument fixes
- The stop gate's block observable is the hook re-firing, not a second action run. A real agent refused at Stop declines to redo finished work; only the mechanical reference driver replays its turn, which is why the bug hid. The matrix was right; the instrument was wrong.
stop_hook_active's value is recorded per invocation, so a re-fire is attributable to the block that caused it.- A
transformtrial where the hook fired and nothing ran measurestransform: false;nullis reserved for the genuinely undecidable shape. evidence_report.diff_against()flagslive-run→live-run-partialas a weakening. A report records which canonicaleventit measured, so apre_toolreport cannot be merged into astopclaim.- Seven tests that could not pass on Windows (POSIX
HOME, cp1252 reads of em dashes, the execute bit) now pass on a clean checkout. None was skipped or deleted.
Docs and CI
- README rewritten to lead with the hero GIF and the honest capability matrix — all 16 agents, a
Verifiedcolumn, and the count ofpre_toolclaims that are live-run witnessed versus doc-derived (4 of 12). The Quick start block runs for real in CI. demo-gifworkflow renders and commits the GIF into the branch that changed the tape;quickstartis its own workflow and diff-checksdocs/quickstart.sh.
Compatibility notes
data/matrix-evidence.jsonis gone;matrix_evidence.EVIDENCEis derived frommatrix.json. Public accessors (matrix.capability,matrix.enforcement_level,matrix_evidence.EVIDENCE) keep their signatures.REVERSE_EVENT_MAPand whatinstallwrites are unchanged for every agent.- Runtime remains stdlib-only.
Full changelog: v0.2.1...v0.3.0