Skip to content

v0.3.0

Choose a tag to compare

@jothimani-rajendran jothimani-rajendran released this 20 Sep 14:40
· 53 commits to main since this release
c84f466

agentseam 0.3.0

The evidence release. Claims on the capability matrix now carry their own evidence records, grades are capped by the kind of evidence behind them, a witnessed run can be frozen and replayed in CI, and Claude Code's three blocking events are witnessed on one real build. Eight adapter defects found by the vendor-truth review are fixed. No wire-format break: every consumer's hook config installs the same.

Upgrading from 0.2.1: pip install agentseam==0.3.0. Consumers that vendor bundles (bundler.bundle(agent)) will see them regenerate — the adapter fixes below are in the bundle, so re-run your bundling step and commit the result.

Highlights

  • Claude Code's prompt_submit and stop cells are witnessed at 2.1.263, the way pre_tool already was. Both gates re-run against the real CLI, twice each, identical results. Two facts the run established: a headless {"decision": "block"} at Stop refuses the agent permission to finish and sends it round again, and Claude Code caps that loop at eight re-fires before ending the turn regardless.
  • Recorded driver — freeze a witnessed run, replay it in CI. One immutable file per witnessed (agent, version) under data/recordings/, read by the package via agentseam.recordings. tools/experiment.py run --record freezes a real-agent run; --driver recorded replays it through the same classifier a live run uses, with no process launched — the seven claude_code@2.1.263 trials reproduce in under a second. The basis chain is now claim → recording → live run.
  • Per-claim evidence, and grading capped by basis. Every asserted cell field (block, rewrite, fail_mode) carries its own {basis, date, version?, test | method} record. matrix.enforcement_level() no longer returns a grade its basis cannot support — a vendor-docs cell asserting fail-closed grades best-effort, never enforced (the defect that shipped as chock#89). Ceiling table: matrix_terms.GRADE_CEILING. No current row's grade changes; the gap is closed defensively.
  • Three new optional cell fields for behaviours the experiment kit already measures: silence_means, timeout_fail_mode, unknown_verb_means. Absence is not a claim. A recognised field a cell does not carry now reads unasserted, never unrecorded by omission.
  • An escalate trial, answered in each agent's own dialect (Cursor spells it ask). New measured field escalate_means, which can read prompted — the run ended waiting on an answer nobody gave.
  • The grade cap honours verified.observed. A live-run-partial row can no longer let an unobserved event back a grade as high as enforced; unobserved events fall back to the row's verified.fallback_basis (default vendor-docs). 11 of 91 claimed pairs change basis; none changes grade.

Adapter fixes (vendor-truth review)

Frozen wire output moves by a handful of bytes in total; each is called out in the changelog.

  • Cursor — the prompt gate no longer surfaces an allow's own rationale in the UI on every submitted prompt, and explains a refusal that was really a degraded rewrite. A Decision.rewrite(None, …) at preToolUse is no longer blamed on the one gate that can express a rewrite. respond() infers an unnamed edits[] payload's event the same way parse() does, so the two can no longer diverge into a permission verdict at a read-only event.
  • Kimi Code — a payload that self-identifies as Kimi but names an event Kimi has not mapped now reaches the caller as UNKNOWN instead of "unrecognized payload" (claims.accept_any_name). A Kimi PermissionRequest is no longer also claimed by Devin (claims.reject_client_types), which had left it unidentified. A control character in an installed command no longer renders a config.toml the vendor refuses to load at all — the full TOML basic-string escape set is emitted and pinned by a tomllib round-trip.
  • Devin — the degraded-rewrite note at UserPromptSubmit/Stop no longer describes a tool call those events don't have.
  • Grok and Kimi Code — PostCompact no longer masquerades as canonical pre_compact; a handler that snapshots context before compaction no longer also fires after it.

Instrument fixes

  • The stop gate's block observable is the hook re-firing, not a second action run. A real agent refused at Stop declines to redo finished work; only the mechanical reference driver replays its turn, which is why the bug hid. The matrix was right; the instrument was wrong.
  • stop_hook_active's value is recorded per invocation, so a re-fire is attributable to the block that caused it.
  • A transform trial where the hook fired and nothing ran measures transform: false; null is reserved for the genuinely undecidable shape.
  • evidence_report.diff_against() flags live-run → live-run-partial as a weakening. A report records which canonical event it measured, so a pre_tool report cannot be merged into a stop claim.
  • Seven tests that could not pass on Windows (POSIX HOME, cp1252 reads of em dashes, the execute bit) now pass on a clean checkout. None was skipped or deleted.

Docs and CI

  • README rewritten to lead with the hero GIF and the honest capability matrix — all 16 agents, a Verified column, and the count of pre_tool claims that are live-run witnessed versus doc-derived (4 of 12). The Quick start block runs for real in CI.
  • demo-gif workflow renders and commits the GIF into the branch that changed the tape; quickstart is its own workflow and diff-checks docs/quickstart.sh.

Compatibility notes

  • data/matrix-evidence.json is gone; matrix_evidence.EVIDENCE is derived from matrix.json. Public accessors (matrix.capability, matrix.enforcement_level, matrix_evidence.EVIDENCE) keep their signatures.
  • REVERSE_EVENT_MAP and what install writes are unchanged for every agent.
  • Runtime remains stdlib-only.

Full changelog: v0.2.1...v0.3.0