v0.4.0 — Agent Security Evaluation Kit
agent-security-bench 0.4.0
Agent Security Evaluation Kit — score agent code and refuse unsafe tools.
Highlights
- Wheel-bundled tasks/security/schema/examples (cold `asb selftest`)
- Frontier tasks ml.16–ml.23 (MCP args, HTML injection, DLP, tenant vectors, cost caps, destructive confirm)
- `asb score-trace` for OpenAI / Anthropic / MCP JSONL
- `asb leaderboard` → task × agent matrix + CI artifact
- Expanded `sec.production_v1` (9 cases)
- Demand map: docs/DEMAND_MAP_2026.md
```bash
pip install -U agent-security-bench==0.4.0
asb selftest
```