Guiderails is a pattern-based steering engine for AI coding agent workflows. It mediates user input, tool calls, and tool output against rule packs. Whether that's a genuinely independent security layer or cooperative enforcement depends on the harness: the external-hook adapter (Claude Code) adds a process boundary the agent can't reach into; the in-process plugin adapter (Pi) doesn't (tamperResistant, see ADR-007).
It is not a security audit tool, a sandbox, or a complete security boundary. Deterministic regex matching cannot catch every adversarial payload. See LIMITATIONS.md for the full statement of what this architecture cannot do, and for the containment pairing (agentjail).
Why redact rather than never-inject? A credential-injection proxy is stronger for secrets registered ahead of time on flows the proxy sees; it doesn't cover unregistered secrets, tokens in files the agent reads, arbitrary command output, or pasted prompts. Never-inject what you know about; redact what you didn't. Pair them (LIMITATIONS.md).
An unmatched tool call allows by default — a coding agent that halts on every unmatched call is unusable. This is a deliberate trade-off, not an oversight (ADR-007). Operators who want a stricter posture can opt in per-category with defaultDecision, or wholesale with the --strict preset.
The exceptions are unconditional and not configurable: an engine crash or timeout always resolves to block on every adapter, and oversized input is rejected by the matcher's MAX_MATCH_INPUT_LENGTH cap rather than silently skipped. Neither is derived from defaultDecision or any user config — there's no legitimate case for allowing a call the engine failed to evaluate.
Depending on your own security model's needs, you could look into pairing it with (non-exhaustive):
- An OS-level sandbox or command-gating tool
- Network-level controls (egress filtering, proxy rules)
- Credential scanning in repositories
- Human review for sensitive operations
- Least-privilege access for the agent's execution environment
See docs/adrs/008-non-goals.md for what this project deliberately does not build.
| Version | Supported |
|---|---|
| 0.1.x | ✅ |
Open a public GitHub issue for any issue with rule packs, matchers, defaults, false negatives, or missing coverage. The reproducer is usually most of the fix, and the community can contribute improvements.
In rare cases where the issue isn't a rule bypass but a fundamental flaw in how the engine itself processes invocations (for example, a condition that silently disables all rule matching), report privately:
- Email: security-guiderails「at」nwo「dot」pm
- GitHub Security Advisory: Use the repository's private advisory feature
- The affected rule pack, matcher, or behavior
- A reproduction case (command, file path, or
ToolCallContextthat should be blocked but isn't) - Expected vs. actual behavior
- Severity assessment (bypass, false negative, false positive, etc.)
- If requesting private disclosure, a short justification against the criteria above
Reports are triaged on a best-effort basis. Reports closed as wontfix will include a rationale, typically noting the issue falls under Known Limitations or is out of scope.
- Regex-based matchers can be evaded via shell tricks (command substitution, string concatenation, alternative tools). The
hardeningrule pack catches many of these via Layer 3 wrapper detection, but not all. - After-tool redaction is the backstop for cases the engine cannot predict. It requires the harness to support the
redactbehavior. - No rule pack covers every secret type or tool. The project ships with common patterns; users should extend rule packs for their specific stack.