-
Notifications
You must be signed in to change notification settings - Fork 0
Policy Engine
aisrf/gateway/policy.py decides, after analysis and before anyone waits, whether a ticket is denied outright, approved without a human, or held for review. It is deterministic, ordered and has no side effects; its output is stored on the ticket as policy_decision and applied by tickets.apply_policy().
evaluate(agent, normalized, path, risk_score, findings) -> PolicyDecision(action, reasons, matched_rules)action is deny, approve or review. Inputs: the Agent row, the normalized request (see Request-Normalization), the stored path (v1/chat/completions style, no leading slash), the integer risk score and the serialised request findings. Global settings are read from the policy namespace of the settings store (settings_store.get_namespace("policy")).
Exactly as in the code; the first rule that returns ends evaluation.
| # | Rule | Driven by | Outcome |
matched_rules entry |
|---|---|---|---|---|
| 1 | Agent disabled | agent.is_active == False |
deny, reason agent is disabled
|
agent.is_active |
| 2 | Global path allow-list |
policy.global_allowed_paths (non-empty list of fnmatch globs); the path must match at least one pattern either as path or as /path
|
deny, reason path '<path>' is not in the global allowed paths
|
policy.global_allowed_paths |
| 3 | Agent path allow-list |
agent.allowed_paths (same matching) |
deny, reason path '<path>' is not in the agent's allowed paths
|
agent.allowed_paths |
| 4 | Agent model allow-list |
agent.allowed_models (fnmatch globs against normalized.model); only evaluated when the request carries a non-empty model |
deny, reason model '<model>' is not in the agent's allowed models
|
agent.allowed_models |
| 5 | Auto-deny regexes |
agent.auto_deny_patterns then policy.global_auto_deny_patterns, each applied with re.search(pat, prompt_text, re.IGNORECASE | re.DOTALL); invalid regexes are skipped |
Every match adds reason prompt matches auto-deny pattern /<pat>/; evaluation continues to rule 6 before returning |
agent.auto_deny_patterns, policy.global_auto_deny_patterns
|
| 6 | Custom rule with deny action | Any finding whose metadata.action == "deny" (produced by the custom rules analyzer from the rules namespace) |
Adds reason a custom rule with action=deny matched; if rules 5 or 6 collected any reason, return deny with all of them |
rules.custom |
| 6b | TypeSafe routing | Only with integrations.typesafe.auto_route; reads the most confident typesafe_guard finding's metadata.typesafe
|
deny with rule typesafe.auto_deny when the verdict is malicious with confidence >= auto_deny_confidence; approve with rule typesafe.auto_approve when it is benign with confidence >= auto_approve_confidence, the agent requires approval and no finding is HIGH or CRITICAL; otherwise continue (TypeSafe-Integration) |
typesafe.auto_deny, typesafe.auto_approve
|
| 7 | Risk auto-deny |
agent.auto_deny_at_risk > 0 and risk_score >= agent.auto_deny_at_risk
|
deny, reason risk score <n> >= auto-deny threshold <t>
|
agent.auto_deny_at_risk |
| 8 | Mandatory review of critical leaks | Any finding with severity == "CRITICAL" and category in (secrets, data_exfil) |
review, reason critical secret/exfiltration finding requires human review; this rule wins over rules 9 and 10 |
finding.critical |
| 9 | Approval not required | agent.require_approval == False |
approve, reason agent policy does not require approval
|
agent.require_approval=false |
| 10 | Risk auto-approve |
agent.auto_approve_below_risk > 0 and risk_score < agent.auto_approve_below_risk
|
approve, reason risk score <n> < auto-approve threshold <t>
|
agent.auto_approve_below_risk |
| 11 | Default | everything else |
review, reason human approval required by agent policy
|
agent.require_approval=true |
Notes:
-
prompt_textin rule 5 is the concatenation of the system prompt and every non-system message (see Request-Normalization); it does not include tool definitions. -
auto_deny_at_riskaccepts 0 to 101 in the API; 0 disables the rule, and since scores are capped at 100, 101 also disables it (the CLI documents101 disables). The default for new agents is 90, which corresponds to theCRITICALband starting at 80 only when a single CRITICAL finding has confidence 1.0; combinations of HIGH findings can reach it too. -
auto_approve_below_riskaccepts 0 to 100; 0 disables. A value of 1 approves only tickets with score 0 (NONE). - The rate limit (
rate_limit_per_minute) and the body size limit are enforced before a ticket exists and are not part of policy. - Rule 8 exists so that a permissive agent (
require_approval=false, or a generous auto-approve threshold) still puts a human in front of a request that appears to carry credentials or an exfiltration payload.
apply_policy() in aisrf/tickets/service.py:
| Action | Ticket status | decided_by |
decision_note |
Gateway behaviour |
|---|---|---|---|---|
deny |
DENIED |
policy |
reasons joined with ;
|
403 with code: denied and risk_score; nothing is sent upstream |
approve |
APPROVED |
policy |
reasons joined with ;
|
Forwarded immediately |
review |
PENDING |
null | empty | Hold registered; sync clients wait, async clients get 202
|
In every case a policy timeline event (actor policy, detail = the decision dict), a per-agent log event policy.<action> and a created SSE event are emitted, and gateway_policy_total{action} is incremented.
Strict production assistant (everything human-reviewed, narrow surface):
{"name": "support-bot", "upstream_provider": "openai", "upstream_api_key": "sk-...",
"require_approval": true, "auto_deny_at_risk": 80,
"allowed_paths": ["v1/chat/completions"], "allowed_models": ["gpt-4o-mini"],
"auto_deny_patterns": ["(?i)ignore (all )?previous instructions", "(?i)drop\\s+table"],
"rate_limit_per_minute": 60}Low-friction internal tool (auto-approve benign traffic, keep humans for anything suspicious):
{"name": "docs-search", "require_approval": true, "auto_approve_below_risk": 25, "auto_deny_at_risk": 90,
"allowed_paths": ["v1/embeddings", "v1/chat/*"]}Tickets scoring 0 to 24 (NONE and LOW) are approved by policy; 25 to 89 wait for a reviewer; 90 and above are denied; a CRITICAL secrets or data_exfil finding always waits regardless of score.
Monitoring only (no gate, full record):
{"name": "batch-summariser", "require_approval": false, "auto_deny_at_risk": 0}Everything is approved and forwarded at once, still ticketed, analysed, logged and audited, except requests with a CRITICAL secrets or exfiltration finding, which are held.
Red-team target with canaries:
{"name": "redteam-target", "require_approval": false, "auto_deny_at_risk": 0, "inject_canary": true}Combine with AISRF_REQUIRE_APPROVAL_FOR_REDTEAM_PROBES to decide whether probes still wait for a human.
CLI equivalent of the first example:
aisrf agent create support-bot --provider openai --upstream-key sk-... \
--auto-deny-at 80 --allowed-path v1/chat/completions --allowed-model gpt-4o-mini \
--deny-pattern '(?i)ignore (all )?previous instructions' --rate-limit 60Stored in app_settings under namespace policy, edited in Settings > Policy or with PUT /api/settings/ns/policy (admin). Defaults from settings_store.NAMESPACE_DEFAULTS:
| Key | Type | Default | Used by |
|---|---|---|---|
global_auto_deny_patterns |
list of regex strings | [] |
Rule 5, after the agent's own patterns |
global_allowed_paths |
list of globs |
[] (all paths) |
Rule 2, before the agent's allow-list |
block_on_canary_leak |
bool | true |
Response withholding (below) |
quarantine_on_critical_response_finding |
bool | false |
Response withholding (below) |
Example:
curl -b jar -X PUT http://localhost:8080/api/settings/ns/policy -H 'Content-Type: application/json' -d '{
"global_allowed_paths": ["v1/chat/*", "v1/messages", "v1/embeddings"],
"global_auto_deny_patterns": ["(?i)reveal (your|the) system prompt"],
"block_on_canary_leak": true,
"quarantine_on_critical_response_finding": true
}'Changes apply live to the next request and persist across restarts. The Settings page includes a browser-side regex tester for the deny patterns.
Policy also runs after the upstream answers, in _finish() of the gateway router (non-streaming responses only; streams are already relayed by the time the body is complete):
- The response text is extracted and the response analyzers run; if the injected canary appears in the text a
CRITICALcanary_leakfinding is inserted first when no analyzer already reported one. - The ticket is marked
COMPLETED(orFAILEDfor status >= 500) with the findings. - If
block_on_canary_leakis true and any response finding has categorycanary_leak: the ticket getserror = "response withheld: system prompt canary leaked in the model output", awithheldtimeline event (actorpolicy),gateway_responses_withheld_totalis incremented, and the client receives403 {"error": {"code": "response_withheld", ...}}instead of the body. - Else if
quarantine_on_critical_response_findingis true and any response finding has severityCRITICAL(for examplesecrets_leak,pii_leak,system_prompt_leakat critical severity): same treatment with the reasonresponse withheld: critical finding in the model output.
The full upstream body remains stored on the ticket (truncated to AISRF_MAX_STORED_RESPONSE_BYTES) so a reviewer can read what was withheld; there is no release action, the client has to resubmit. pipeline.submit() does not withhold; it returns the response findings to the caller (the red-team evaluators use them).
Regex rules from the rules namespace run as an analyzer and produce findings; a rule with action: deny adds metadata.action = "deny" to its finding, which rule 6 turns into a denial. Guardrail integrations (LLM Guard, NeMo, Lakera, Rebuff) also surface as findings and therefore influence the score used by rules 7 and 10. See Custom-Rules and Guardrails.
Related: Core-Concepts, Agents-and-Credentials, Configuration-Reference, Canary-Words.
AISRF, AI Security & Research Framework. github.com/keyuraghao/aisrf, Apache License 2.0.
Start
Gateway
- Gateway-Endpoints-and-Headers
- Request-Normalization
- Policy-Engine
- Agents-and-Credentials
- Configuration-Reference
- Settings-Center
Review
Security analysis
Red teaming
Code review
Interfaces
Operations
Project