Skip to content

Security and Safety

Mohammed Danish Amber edited this page Oct 5, 2026 · 1 revision

Security and Safety

Attacking agents is only acceptable if nothing real is harmed. Safety is built into the engine, not a setting.

Non-negotiable invariants

  • Record-only, fail-closed mirror. Mutating tool calls are captured but never executed; read-only calls pass through; unclassifiable calls are blocked. The run loop passes a no-op executor, so even a mutating call does nothing real.
  • Proof, not judgement. A success requires an actual per-run canary token hit — in the reply or a captured tool call's args — never an LLM's opinion. (metered is a cost threshold, not a canary; see How It Works.)
  • Per-run canaries. Tokens are HMAC-SHA256(run_key, …) with a random run_key per run, so they are unpredictable and worthless if leaked.
  • Sandbox-inert fixture. The bundled target grants no real primitive — see The Fixture.
  • Append-only evidence. Every step is recorded immutably — see Evidence Schema.

These are covered by tests (sandbox-inert checks, mirror routing, verdict boundaries) that run in CI — see Development and Release.

Using it responsibly

  • Only test agents you are authorized to test.
  • Never expose the fixture to a network.
  • Against a non-fixture target, delivered payloads are real attack text — run only against your own systems in a controlled environment.

Reporting a vulnerability in Aphasia Agentry itself

The deliberately-vulnerable fixture is not a vulnerability. A real issue is one in the engine (e.g. the mirror executing a mutating call, a sandbox guarantee failing, a forgeable canary). Report it privately via the repository's Security Advisories — see SECURITY.md in the repo. Do not open a public issue for a security report.

Clone this wiki locally