The first release you can actually install.
pip install agentmetry
agentmetry doctor0.3.0 would have written core, api and cli into site-packages, so it was never published. This one works: doctor opens with no failures on a clean install, and the detection corpus and policy manifests ship inside the package.
Check the detection claims before trusting them, no clone required:
agentmetry benchmarkcases 46 (24 attack, 22 benign)
rules covered 13
detected 24 / 24
missed 0
false positives 0
New
Merkle inclusion proofs. The hash chain answers "has this file been altered". It cannot answer "did this event happen" without handing over the whole file. agentmetry prove <trail> --seq N emits an RFC 6962 proof for one event: 1.2 KB against an 8.3 MB trail. Proving one tool call no longer means disclosing a month of unrelated work.
CloudEvents v1.0 export. AGENTMETRY_AUDIT_WEBHOOK_FORMAT=cloudevents for Knative, EventBridge, Event Grid, Dapr or Kafka. The canonical event still travels whole in data.
Ingest Microsoft Agent Governance Toolkit audit files. agentmetry import-agt verifies AGT's hash chain and HMAC signatures, then runs sequence detection over them. On a file produced by AGT's own FileAuditSink, three individually-permitted calls raised one critical credential-exfil — which is the difference between a per-call gate and a session recorder, rather than a criticism of AGT.
Fixed: the detection engine had holes
Four evasions, none exotic, each under a minute to build, all found by auditing the engine rather than by a detection firing:
| was | |
|---|---|
echo $AWS_SECRET_ACCESS_KEY |
generic Execution |
python -c "urllib.request.urlopen(...)" |
no egress tag |
bash <(curl -fsSL ...) |
no cradle match (no pipe) |
cp -r ~/.ssh /tmp/k |
no traits at all |
Plus credential-exfil could not fire when the read and the send were one command, and remote-staging-then-execute knew exactly seven staging hosts.
Underneath sat one defect: credential recognition existed twice, and the rules read only the MITRE tag — the classifier with less information and no test corpus. A bare .env in its pattern list tagged the module path agentmetry.core.diagnostics.env_file as credential access and manufactured the credential half of two critical findings.
Autostart was registered and failing. The scheduled task launched a module the package rename had removed, exiting 1 every sixty seconds while doctor reported OK and events piled up in the spool. doctor now fails a registration that does not work.
Corpus: 20 cases to 46
The benign half went from 6 sessions to 22. Six benign sessions put a 95% confidence bound on the false-positive rate at roughly 41%, which is not a rate. It is now about 19%.
Every new attack case is paired with the near-miss that must stay silent: autonomous writes before an approval against the same writes after one, fetch-then-egress against fetch-then-edit, downloading a file and running it against downloading a schema and running a repo script. A threshold that drifts now breaks a benign case instead of surfacing quietly in production.
That number is a regression guard, not a field false-positive rate. The field rate is what the four-week dogfood run reports, over traffic nobody chose.
Still true
This is a public alpha. It is a recorder, not a sandbox: the only enforcement path is pre-execution DLP blocking in the hook, and it ships set to log. It is not a CASB — it records the agents you wire in, and an unmanaged browser assistant is invisible to it. The DLP is regex, not ML.
Full detail in CHANGELOG.md.
Published from a tagged CI run via PyPI Trusted Publishing, so the artifact is built from a clean checkout of this tag.