-
Notifications
You must be signed in to change notification settings - Fork 0
Guardrail Pipelines
Guardrail Pipelines add a configurable policy-enforcement layer that intercepts requests before provider routing (pre-call) and responses before returning to the caller (post-call). Use them for PII/DLP filtering, content moderation, semantic prompt guarding, and sensitive data redaction with transparent re-injection.
Client Request
│
▼
┌─────────────────┐
│ PRE-CALL │ ← PII redaction, secret scanning, injection detection
│ Pipeline │
└────────┬────────┘
│ (redacted request)
▼
┌─────────────────┐
│ Router / │ ← Normal provider routing with failover
│ Provider │
└────────┬────────┘
│ (raw response)
▼
┌─────────────────┐
│ POST-CALL │ ← Output filtering, PII re-injection
│ Pipeline │
└────────┬────────┘
│ (clean response)
▼
Client Response
| Type | Description |
|---|---|
regex |
Up to 256 named patterns with allow/deny modes; compiled at load time, per-pattern 10ms budget |
presidio |
Presidio-compatible NLP PII detection via HTTP; configurable entity types and confidence threshold |
openai_moderation |
OpenAI Moderation API integration |
lakera |
Lakera Guard prompt injection detection |
semantic |
Embedding-based similarity matching against allow/deny example collections in Qdrant |
custom_http |
POST content to any HTTP endpoint implementing the documented findings JSON schema |
| Action | Pre-Call | Post-Call | Behavior |
|---|---|---|---|
allow |
✓ | ✓ | Pass through unmodified |
block |
✓ | ✓ | Reject with HTTP 403 |
mask |
✓ | Replace each character with *, preserving byte length |
|
redact |
✓ | ✓ | Replace with placeholder tokens (pre-call) or [REDACTED] (post-call) |
replace_with_policy_message |
✓ | Replace assistant content with a configured message |
guardrails:
max_reinjection_entries: 256 # PII values tracked for re-injection (1–10000; default 256)
providers:
- name: secret-scanner
type: regex
failure_policy: fail_close
patterns:
- { name: openai_key, regex: "sk-[A-Za-z0-9]{20,}", entity: API_KEY, mode: deny }
- { name: aws_key, regex: "AKIA[0-9A-Z]{16}", entity: AWS_KEY, mode: deny }
- name: pii-detector
type: presidio
failure_policy: fail_open
endpoint: "http://presidio:3000/analyze"
language: en
entities: [EMAIL_ADDRESS, US_SSN, CREDIT_CARD]
confidence_threshold: 0.6
- name: prompt-guard
type: semantic
failure_policy: fail_open
allow_collection: "guardrail_allow"
deny_collection: "guardrail_deny"
allow_threshold: 0.90
deny_threshold: 0.85
pipelines:
- name: standard
failover_on_refusal: true
redaction_notice_instruction: null # null = use DEFAULT_REDACTION_NOTICE_INSTRUCTION
instruction_insertion_mode: separate # separate | merged
refusal_phrase_list: null # null = use DEFAULT_REFUSAL_PHRASES
stages:
- { name: pii-redact, provider: pii-detector, phase: pre_call, action: redact }
- { name: secret-block, provider: secret-scanner, phase: pre_call, action: block }
- { name: injection-guard, provider: prompt-guard, phase: pre_call, action: block }
- { name: out-redact, provider: secret-scanner, phase: post_call, action: redact }
global_default_pipeline: standard
bindings:
virtual_keys:
vk_team_a: standard
model_groups:
gpt-4-group: standard
routes:
"/v1/chat/completions": standard
failover_on_refusal:
vk_external: trueThe Guardrails tab in the admin panel provides two modes:
Quick Setup — Choose a common protection preset or create a rule in plain language:
| Preset | Effect |
|---|---|
| Block harmful content | Adds OpenAI Moderation as a pre-call block stage |
| Protect personal information | Adds PII detection with pre-call redact action |
| Block prompt attacks | Adds Lakera Guard as a pre-call block stage |
| Try another model after a refusal | Enables failover_on_refusal on the pipeline |
Create a custom rule — Select a rule name, detection type (harmful/unsafe content, personal information, prompt injection, custom regex), phase (before/after sending to the model), and action (block, redact, mask, allow).
Advanced setup — Expandable section that exposes the full YAML-equivalent provider, pipeline, and binding configuration for fine-grained control.

When a pre-call stage uses the redact action, detected PII is replaced with deterministic placeholder tokens before the request reaches the LLM:
User: "Contact me at john@example.com or 555-0123"
↓ (redacted)
LLM sees: "Contact me at <<PII_EMAIL_1>> or <<PII_PHONE_1>>"
↓ (LLM responds)
LLM output: "I'll reach out to <<PII_EMAIL_1>>"
↓ (re-injected)
Client receives: "I'll reach out to john@example.com"
- PII values are replaced with
<<PII_TYPE_N>>placeholders - A system instruction is prepended telling the model to preserve placeholders verbatim
- After the LLM responds, placeholders are restored to original values
- Up to
max_reinjection_entriesdistinct values per request receive re-injection entries (configurable 1–10000; default 256) - Identical values reuse the same placeholder (deduplication)
- Configurable redaction-notice instruction text and insertion mode (
separateormerged) - The re-injection map is held only in memory for the request duration
- Values beyond the cap are still redacted but placeholders are not recorded for restoration
When PII is redacted, the gateway prepends an instruction to the request telling the model to preserve placeholders. This instruction is configurable per-pipeline:
pipelines:
- name: standard
redaction_notice_instruction: null # null = use built-in default
instruction_insertion_mode: separate # separate | merged| Mode | Behavior |
|---|---|
separate (default) |
Inserts a new system message before all existing messages |
merged |
Prepends the instruction into the existing leading system message (with blank-line separator) |
The built-in default instruction tells the model:
- Treat placeholder tokens as valid input (don't refuse or warn)
- Proceed normally with the task, using placeholders where real values would appear
- Reproduce placeholders verbatim so the downstream layer can restore originals
The gateway detects when a model refuses a request and optionally fails over:
| Signal | Description |
|---|---|
| Phrase matching | Case-insensitive regex patterns against assistant content (default list + configurable) |
| Tool-omission | Tools were provided but the model didn't call any |
- Re-dispatches the already-redacted request to the next eligible target
- Skips providers with open circuit breakers
- Each provider attempted at most once
- Toggle:
failover_on_refusalper-pipeline or per-binding (disabled by default)
Each pipeline can override the default refusal phrase list:
pipelines:
- name: strict
failover_on_refusal: true
refusal_phrase_list:
- "I cannot"
- "I'm unable to"
- "against my guidelines"
- "I must decline"When refusal_phrase_list is absent (null), the built-in DEFAULT_REFUSAL_PHRASES list is used. Each entry is a case-insensitive regex pattern — validation rejects empty or uncompilable entries.
Bindings can also enable refusal-failover independently:
bindings:
failover_on_refusal:
vk_external: true # Enables failover for this virtual key regardless of pipeline settingWhen multiple pipelines apply (global + virtual-key + model-group + route), stages concatenate in a fixed order:
- Global default pipeline stages
- Virtual-key pipeline stages
- Model-group pipeline stages
- Route pipeline stages
Halting actions (block, replace_with_policy_message) short-circuit immediately. Non-halting actions continue to the next stage.
Each provider must declare a failure_policy:
| Policy | Behavior |
|---|---|
fail_open |
On timeout or error, skip the stage and continue |
fail_close |
On timeout or error, halt pipeline and return HTTP 503 |
For SSE responses with a post-call pipeline:
- Gateway buffers the streamed response (up to 10 MB)
- Sends keep-alive comments during buffering
- Applies guardrail analysis on the assembled content
- Re-chunks the result into SSE events matching original chunk boundaries
Guardrail execution is fully observable via Prometheus:
| Metric | Description |
|---|---|
obey_api_guardrail_stage_executions_total{pipeline, stage, provider, action} |
Stage execution counter |
obey_api_guardrail_stage_latency_ms{pipeline, stage, provider} |
Latency histogram (buckets: 5–5000ms) |
obey_api_guardrail_refusal_detected_total{pipeline, signal} |
Refusal detection counter |
- Virtual Keys — bind pipelines to specific callers
- Security — encryption and secrets management
- Providers — provider configuration