Skip to content

Releases: kunko-ai-labs/agent-assurance

v0.6.0

Choose a tag to compare

@github-actions github-actions released this 21 Sep 14:18
9607446

What's Changed

  • chore(github): issue forms, PR template, CODEOWNERS by @vmbb13 in #33
  • ci: fork-safe workflows (no signing / no comment with a read-only token) by @vmbb13 in #34
  • chore(github): [EP-000] issue templates (Epic, User Story, Bug, Catalogue), structured labels, PR convention by @vmbb13 in #35
  • feat(cli): [US-001-003] validate policy files and report name, sha256 and overrides by @makiaveli1 in #31
  • feat(mcp): [EP-001] Declare read-only annotations on every MCP tool by @vmbb13 in #40
  • release: v0.6.0 — MCP tool annotations, policy validation, fork-safe CI, formal templates by @vmbb13 in #41

New Contributors

Full Changelog: v0.5.2...v0.6.0

v0.5.2 — Discoverable on PyPI

Choose a tag to compare

@github-actions github-actions released this 18 Sep 09:54
c03c4a9

`pipx install agent-assurance` — every release ships Sigstore-signed SLSA provenance (`gh attestation verify agent_assurance-0.5.2-py3-none-any.whl -R kunko-ai-labs/agent-assurance`).

  • PyPI metadata: trove classifiers (audience, licence, Python 3.10–3.12, Security / QA) and project URLs (docs, changelog, issues, security policy)
  • README: the SHA-pinning guidance shows how to fetch a release's commit instead of a placeholder that some renderers swallowed; version mentions unified
  • The release pipeline now creates the GitHub release itself when a tag arrives without one

Contract unchanged: exit `0` pass · `1` gate · `2` usage.

v0.5.1 — First PyPI release

Choose a tag to compare

@vmbb13 vmbb13 released this 18 Sep 08:39
1e9ce18

First release on PyPI: pipx install agent-assurance (Trusted Publishing, no stored tokens). The wheel and sdist attached here carry Sigstore build provenance: gh attestation verify agent_assurance-0.5.1-py3-none-any.whl -R kunko-ai-labs/agent-assurance.

Since v0.5.0: repo public and on the GitHub Marketplace; every third-party action pinned by SHA; Dependabot, OpenSSF Scorecard and CodeQL; README trust checklist (SHA pinning, no network calls, least privilege), comparison table, FAQ; SECURITY.md; release pipeline with build provenance.

Contract unchanged: exit 0 pass · 1 gate · 2 usage.

v0.5.0 — Readable outside the repo, tunable without forking

Choose a tag to compare

@vmbb13 vmbb13 released this 17 Sep 17:02

What's new

  • Capability card — --format html (scan, check, diff): a single self-contained file, no scripts, no network. Verdict, promise, observed capabilities with file:line and human-in-the-loop, score breakdown, sources, policy, tool version. Action input card: uploads it as an artifact.
  • Organisation policy — agent-assurance.policy.yaml: weights, bands, gate, which classes break a promise, extra read-only commands, your own MCP servers (no longer UNKNOWN), extra tool-definition files. Defaults equal the built-in model exactly (tested). Every report/card/attestation records the policy name and sha256. JSON schema included.
  • Codex CLI — .codex/config.toml [mcp_servers.<name>] shares the MCP classification.
  • Serialized tool definitions — agent-tools.json|yaml (OpenAI, Agents SDK, MCP, Claude, LangChain/CrewAI shapes). Class inferred from verbs and marked inferred: scores like its class, but never breaks a promise on its own (review). Nothing imported or executed.

Contract unchanged: exit 0 pass · 1 gate · 2 usage.

v0.4.0 — Evidence that survives the repo

Choose a tag to compare

@vmbb13 vmbb13 released this 17 Sep 16:45

What's new

  • agent-assurance attest <dir> -o aa-attestation.json — an in-toto Statement v1: the manifest and every parsed configuration file as subjects (sha256); predicate = tool version, timestamp, git commit, declared promise, observed capabilities, full report, standards touched. Unsigned, it is an archivable audit record. Exit code still reports the gate, so a pipeline can keep the evidence and stop on a broken promise.
  • Action input attest — write uploads the JSON as an artifact; sign hands it to actions/attest (Sigstore, bound to the workflow identity; needs id-token: write, attestations: write, and a public repo or GitHub Enterprise). Verify later with gh attestation verify <file> -R org/repo --predicate-type https://github.com/kunko-ai-labs/agent-assurance/attestation/v1.
  • docs/real-world.md — four public repositories scanned as-is, pinned to commit.
  • README: real examples of what it catches, the live PR comment, the evidence section.

Mapping to EU AI Act art. 12 and SOC 2 change management is adapted, not a conformance claim: this is the capabilities half of the record; runtime logs are the other half.

v0.3.0 — One promise, three enforcement points

Choose a tag to compare

@vmbb13 vmbb13 released this 17 Sep 16:33

What's new

  • agent-assurance diff <base> <head> — what did this change do to the agent's reach and promise? Capabilities added/removed, band movement, promise kept → broken with the exact newly broken items. Cosmetic changes produce no delta. --fail-on-delta gates on regression. SARIF carries baselineState.
  • Action mode: diff — on pull_request, checks out the base into a detached worktree, runs the diff, fails the job and posts one PR comment that is updated on every push. See the open demo PR, red on purpose.
  • MCP server (pip install "agent-assurance[mcp]", agent-assurance-mcp) — tools scan, check, would_break: the agent asks "if I add this server, does the promise break?" before touching a file. Simulated in a temp copy; nothing written.
  • Claude Code PostToolUse hook (contrib/claude-code/) — the agent is told, in the same turn, that its edit broke the promise.
  • One MCP scanner for four hosts: .mcp.json, .cursor/mcp.json, .gemini/settings.json, .vscode/mcp.json. Same server declared for several hosts counts once.
  • pre-commit hook (agent-assurance-scan).
  • README rewritten; demo GIF.

Fixed after scanning eight real public repos

  • Ten allow: Bash(...) rules were scored as ten unattended shells (a benign repo came out CRITICAL for grep, git status, npm test). Rules now collapse per capability class; read-only commands are read; scoped grants are not penalised and AA-002 reports them as review, not broken.
  • aws-knowledge (docs, read-only) matched the aws entry (delete). Catalogue: aws-docs, context7, todoist added; matching stays package-first, token-based on names.
  • Claude Code housekeeping tools (Skill, TodoWrite, …) were UNKNOWN.
  • A step name containing : broke action.yml; now quoted and covered by a YAML shape test.

Contract unchanged: exit 0 pass · 1 gate · 2 usage.

v0.2.0 — Declare what your agent may do. Verify it on every PR.

Choose a tag to compare

@vmbb13 vmbb13 released this 17 Sep 16:03

What's new

agent-assurance scan <dir> observes what a repo's agent configuration actually grants and verifies it against the declared manifest.

  • AA-002 Declared vs Observed — the manifest is a promise; the scan is the evidence. Undeclared write / delete / execute / external_send / financial capability or sensitive data → FAIL, anchored at the file:line that broke it (markdown, JSON, SARIF). Unknown → never a silent PASS.
  • .mcp.json — servers resolved through a curated, sourced catalogue of MCP servers; remote url = egress; credential-looking env/header names = credential access (values never read). Unknown servers are UNKNOWN.
  • .claude/settings.json — Claude Code permission rules. A rule does not grant a capability, it decides whether a human approves it: allow = auto, ask = person, deny = removed. Declared autonomy: 2 + allow: Bash(*) = promise broken.
  • Sources scanned table in every report: parsed, detected-but-unsupported, what was deduced.
  • Action mode: scan with directory; dogfood CI asserts PASS on kept promises and BLOCK on broken ones.

Fixed since v0.1.0

  • SARIF pointed at a fixed agent-assurance.yaml and lacked a region; now anchors at the real file and line, PASS results are kind: pass (no Security-tab noise).
  • Exit codes are a tested contract: 0 pass · 1 gate · 2 usage.
  • Dogfood CI was a demo (continue-on-error without assertion); it now fails if the gate stops blocking.

Why this and not another permission-diff bot

See docs/landscape.md: vulnerability scanners find poisoned tools; permission-diff bots show what changed; none verify a declared promise or map to OWASP Agentic Top 10. This does both, deterministically, with no LLM in the verdict.

Full contract and roadmap in the README.