v0.2.0
agent-security-bench v0.2.0
Open bench for evaluating AI coding agents on ML tasks and agent security, with machine-readable receipts.
Included
- 10 ML tasks (
tasks/ml/) - Core + extended security suites (prompt injection, jailbreak, tool overreach)
- CLI:
list,score,security,batch,validate-receipt - MCP fixture server
- Examples, CI, receipt schema
Install (after PyPI publish finishes)
pip install agent-security-bench==0.2.0Related: https://github.com/sinakazemnezhad/persian-llm-reference