Skip to content

Releases: open-coder-ai/agentseam

v0.3.5

Choose a tag to compare

@jothimani-rajendran jothimani-rajendran released this 29 Sep 01:58
c5d2a00

What's Changed

Full Changelog: v0.3.4...v0.3.5

v0.3.4

Choose a tag to compare

@jothimani-rajendran jothimani-rajendran released this 27 Sep 19:04
d3e083e

What's Changed

Full Changelog: v0.3.3...v0.3.4

v0.3.3

Choose a tag to compare

@jothimani-rajendran jothimani-rajendran released this 23 Sep 10:41
a7e712b

What's Changed

Full Changelog: v0.3.2...v0.3.3

v0.3.0

Choose a tag to compare

@jothimani-rajendran jothimani-rajendran released this 20 Sep 14:40
c84f466

agentseam 0.3.0

The evidence release. Claims on the capability matrix now carry their own evidence records, grades are capped by the kind of evidence behind them, a witnessed run can be frozen and replayed in CI, and Claude Code's three blocking events are witnessed on one real build. Eight adapter defects found by the vendor-truth review are fixed. No wire-format break: every consumer's hook config installs the same.

Upgrading from 0.2.1: pip install agentseam==0.3.0. Consumers that vendor bundles (bundler.bundle(agent)) will see them regenerate — the adapter fixes below are in the bundle, so re-run your bundling step and commit the result.

Highlights

  • Claude Code's prompt_submit and stop cells are witnessed at 2.1.263, the way pre_tool already was. Both gates re-run against the real CLI, twice each, identical results. Two facts the run established: a headless {"decision": "block"} at Stop refuses the agent permission to finish and sends it round again, and Claude Code caps that loop at eight re-fires before ending the turn regardless.
  • Recorded driver — freeze a witnessed run, replay it in CI. One immutable file per witnessed (agent, version) under data/recordings/, read by the package via agentseam.recordings. tools/experiment.py run --record freezes a real-agent run; --driver recorded replays it through the same classifier a live run uses, with no process launched — the seven claude_code@2.1.263 trials reproduce in under a second. The basis chain is now claim → recording → live run.
  • Per-claim evidence, and grading capped by basis. Every asserted cell field (block, rewrite, fail_mode) carries its own {basis, date, version?, test | method} record. matrix.enforcement_level() no longer returns a grade its basis cannot support — a vendor-docs cell asserting fail-closed grades best-effort, never enforced (the defect that shipped as chock#89). Ceiling table: matrix_terms.GRADE_CEILING. No current row's grade changes; the gap is closed defensively.
  • Three new optional cell fields for behaviours the experiment kit already measures: silence_means, timeout_fail_mode, unknown_verb_means. Absence is not a claim. A recognised field a cell does not carry now reads unasserted, never unrecorded by omission.
  • An escalate trial, answered in each agent's own dialect (Cursor spells it ask). New measured field escalate_means, which can read prompted — the run ended waiting on an answer nobody gave.
  • The grade cap honours verified.observed. A live-run-partial row can no longer let an unobserved event back a grade as high as enforced; unobserved events fall back to the row's verified.fallback_basis (default vendor-docs). 11 of 91 claimed pairs change basis; none changes grade.

Adapter fixes (vendor-truth review)

Frozen wire output moves by a handful of bytes in total; each is called out in the changelog.

  • Cursor — the prompt gate no longer surfaces an allow's own rationale in the UI on every submitted prompt, and explains a refusal that was really a degraded rewrite. A Decision.rewrite(None, …) at preToolUse is no longer blamed on the one gate that can express a rewrite. respond() infers an unnamed edits[] payload's event the same way parse() does, so the two can no longer diverge into a permission verdict at a read-only event.
  • Kimi Code — a payload that self-identifies as Kimi but names an event Kimi has not mapped now reaches the caller as UNKNOWN instead of "unrecognized payload" (claims.accept_any_name). A Kimi PermissionRequest is no longer also claimed by Devin (claims.reject_client_types), which had left it unidentified. A control character in an installed command no longer renders a config.toml the vendor refuses to load at all — the full TOML basic-string escape set is emitted and pinned by a tomllib round-trip.
  • Devin — the degraded-rewrite note at UserPromptSubmit/Stop no longer describes a tool call those events don't have.
  • Grok and Kimi Code — PostCompact no longer masquerades as canonical pre_compact; a handler that snapshots context before compaction no longer also fires after it.

Instrument fixes

  • The stop gate's block observable is the hook re-firing, not a second action run. A real agent refused at Stop declines to redo finished work; only the mechanical reference driver replays its turn, which is why the bug hid. The matrix was right; the instrument was wrong.
  • stop_hook_active's value is recorded per invocation, so a re-fire is attributable to the block that caused it.
  • A transform trial where the hook fired and nothing ran measures transform: false; null is reserved for the genuinely undecidable shape.
  • evidence_report.diff_against() flags live-run → live-run-partial as a weakening. A report records which canonical event it measured, so a pre_tool report cannot be merged into a stop claim.
  • Seven tests that could not pass on Windows (POSIX HOME, cp1252 reads of em dashes, the execute bit) now pass on a clean checkout. None was skipped or deleted.

Docs and CI

  • README rewritten to lead with the hero GIF and the honest capability matrix — all 16 agents, a Verified column, and the count of pre_tool claims that are live-run witnessed versus doc-derived (4 of 12). The Quick start block runs for real in CI.
  • demo-gif workflow renders and commits the GIF into the branch that changed the tape; quickstart is its own workflow and diff-checks docs/quickstart.sh.

Compatibility notes

  • data/matrix-evidence.json is gone; matrix_evidence.EVIDENCE is derived from matrix.json. Public accessors (matrix.capability, matrix.enforcement_level, matrix_evidence.EVIDENCE) keep their signatures.
  • REVERSE_EVENT_MAP and what install writes are unchanged for every agent.
  • Runtime remains stdlib-only.

Full changelog: v0.2.1...v0.3.0

v0.2.1

Choose a tag to compare

@jothimani-rajendran jothimani-rajendran released this 02 Sep 10:14
7337d4d

A patch release: three waves of internal hardening since 0.2.0, no wire or behaviour
change anywhere — every golden fixture stays byte-identical.

Closed the vendor-config gaps C2 found

Executing chock's C2 wave against the 0.2.0 wheel surfaced three real gaps in the
vendor data. Two are now recorded with full evidence: tools.shell for
vscode_copilot and codex_cli (each vendor's own documented tool-name vocabulary),
and repo_root_token, an opt-in field for a vendor's project-directory placeholder,
populated only where a primary source documents one. The third — Copilot's native
hooks envelope, a second vendor-documented config surface distinct from the shape this
adapter emits — has no honest home in the current schema without new constructs, so
it's recorded as design feedback rather than forced into data that would overclaim.

Externalized the last stranded prose and templates

Following the externalize, don't hardcode standard: the bare-ALLOW audit, packaging
limits, and per-vendor content-rule reasons move from Python string literals into
schema-validated JSON tables. The bundle header, runtime, and vendor-binding source
move to .py.tmpl template files with __TOKEN__ placeholders, so each one is valid
Python as committed and lints as such — no more logic hiding inside multi-line strings.

Adopted a Sonar/Checkstyle/FindBugs-class lint bar

Turned on ruff's complexity, naming, boolean-trap, exception-hygiene, dead-code, and
CLI-print-discipline rule families across src/, plus a new literal-duplication guard.
Every finding fixed category-by-category in bisectable commits — pure refactor, no
behavioural change. tests/, tools/, examples/, and two docs/ asset scripts are
scoped out for a follow-up wave, each with a TODO(lint-adoption) reason rather than a
blanket exemption.

Verification

Full test suite green throughout (1448 passed), golden fixture replay, bundle
equivalence, and the 12-agent subprocess replay all pass unchanged.

Full changelog: https://github.com/open-coder-ai/agentseam/blob/main/CHANGELOG.md#021---2026-09-02

v0.2.0

Choose a tag to compare

@jothimani-rajendran jothimani-rajendran released this 01 Sep 21:56
acb3881

agentseam 0.2.0 rebuilds how vendor support is expressed: eleven of the twelve
adapted agents are no longer Python modules but JSON config entries driving a
shared family engine. Adding a comparable vendor is now a config file, not code —
and that is a tested property, not a claim: the suite adds a synthetic
thirteenth vendor by config alone and proves its bundle matches the library.

Highlights

  • Vendor config layer (src/agentseam/data/vendors/): a schema plus one
    entry per agent, carrying events, field chains, claims markers, verdict
    dialects, hook-entry shapes — and per-claim evidence (basis/date/test);
    schema validation fails any claim without a basis. Every entry is derived by
    a committed recount tool that executes or AST-reads the source of truth,
    never hand-transcribed.
  • Family engines replace bespoke adapters for claude_code, codex_cli,
    kimi_code, devin, gemini_cli, tabnine, junie, grok, cursor, windsurf and
    antigravity. vscode_copilot remains a dialect module by recorded design
    decision (its claims/parse logic branches on payload content).
  • Bundles are engine + entry, trimmed per family: each single-file plugin
    now carries only the renderer grammars its gates speak, plus its own vendor
    literal, with a structural test enforcing exactly one of each.

Changed / deprecated

  • Decision vocabulary aligned to the Agent Control Standard: ASK →
    ESCALATE, REWRITE → TRANSFORM. The old names and
    Decision.ask()/Decision.rewrite() remain as deprecated aliases (the
    classmethods emit DeprecationWarning). New WARN outcome; no adapted
    agent has an established warning channel yet, so it degrades to an
    honestly-labelled allow.
  • matrix.capability() returns both "rewrite" and "transform" keys (same
    value) for one minor version; can_transform() joins can_rewrite().
  • Breaking: degraded_from in a decision's evidence now reports the
    canonical spelling ("transform", not "rewrite").
  • Bundle output bytes differ from 0.1.1 (comment strip + recomposition);
    regenerate vendored bundles with this version rather than diffing against
    0.1.1 output.

What did not change

The wire. Every payload → response pair an agent sees is byte-for-byte
identical to 0.1.1, enforced by a golden fixture suite (546 frozen wire
exchanges across all 12 agents) that every change in this release was gated
on and which was never regenerated.

Fixed

  • tools/capture.py probe file permissions tightened to 0o700 (CodeQL
    py/overly-permissive-file).
  • matrix.json's kimi_code row named the wrong config file; corrected via
    recount, evidence record untouched.
  • pyproject.toml's version is now test-checked against
    agentseam.__version__ (a wheel can no longer ship disagreeing).
  • CI installs from hash-pinned lockfiles; OSSF Scorecard Pinned-Dependencies
    4 → 10.

Full detail: CHANGELOG.md.