48-scenario benchmark testing identity, policy enforcement, and observability across AI agent runtimes.
-
Updated
Jul 13, 2026 - Python
48-scenario benchmark testing identity, policy enforcement, and observability across AI agent runtimes.
TypeScript + Python SDKs for scope-narrowed subagents and delegation chains — every action audits back to the originating human. Anthropic SDK, CrewAI, LangChain.
Add a description, image, and links to the tool-permissions topic page so that developers can more easily learn about it.
To associate your repository with the tool-permissions topic, visit your repo's landing page and select "manage topics."