-
Notifications
You must be signed in to change notification settings - Fork 0
Security and Safety
Mohammed Danish Amber edited this page Oct 5, 2026
·
1 revision
Attacking agents is only acceptable if nothing real is harmed. Safety is built into the engine, not a setting.
- Record-only, fail-closed mirror. Mutating tool calls are captured but never executed; read-only calls pass through; unclassifiable calls are blocked. The run loop passes a no-op executor, so even a mutating call does nothing real.
-
Proof, not judgement. A
successrequires an actual per-run canary token hit — in the reply or a captured tool call's args — never an LLM's opinion. (meteredis a cost threshold, not a canary; see How It Works.) -
Per-run canaries. Tokens are
HMAC-SHA256(run_key, …)with a randomrun_keyper run, so they are unpredictable and worthless if leaked. - Sandbox-inert fixture. The bundled target grants no real primitive — see The Fixture.
- Append-only evidence. Every step is recorded immutably — see Evidence Schema.
These are covered by tests (sandbox-inert checks, mirror routing, verdict boundaries) that run in CI — see Development and Release.
- Only test agents you are authorized to test.
- Never expose the fixture to a network.
- Against a non-fixture target, delivered payloads are real attack text — run only against your own systems in a controlled environment.
The deliberately-vulnerable fixture is not a vulnerability. A real issue is one in the
engine (e.g. the mirror executing a mutating call, a sandbox guarantee failing, a forgeable
canary). Report it privately via the repository's Security Advisories — see SECURITY.md
in the repo. Do not open a public issue for a security report.