Skip to content

10. Decision Storage and Model Fallback

Hemant Kohli edited this page Sep 23, 2026 · 2 revisions

Persistent decisions and fallback guard models

Configure persistent audit delivery and explicit fallback for the local intent and PII models. It does not replace privacy-filter or introduce generative-model fallback.

Shared configuration

import { defineGuardConfig, createJsonlDecisionStore } from "@intflows/genkit-guard";

export const decisions = createJsonlDecisionStore("./logs/guard-decisions.jsonl");
export default defineGuardConfig({
  policyVersion: "support-v2",
  models: {
    extractor: "Xenova/all-MiniLM-L6-v2",
    // Replace with a prepared, compatible feature-extraction model:
    extractorFallback: "your-org/intent-backup"
  },
  intent: { semantic: {
    threshold: 0.7,
    intents: { support: "Technical support for Azure Blob Storage and APIs" }
  } },
  pii: {
    mode: "classifier",
    model: "openai/privacy-filter",
    fallback: {
      mode: "ner",
      model: "Xenova/bert-base-NER",
      labelMappings: { PER: "NAME" }
    }
  },
  tools: { defaultAction: "block", rules: { lookupTicket: "redact" } },
  logging: { enabled: false, store: decisions }
});

Pass the same configuration to initGuard and guard. No store or fallback is enabled by default. Tool rules apply to separately defined Genkit tools.

Persistent delivery

GuardDecisionStore has one method: append(decision): void | Promise<void>. Use createGuardDecisionStore({ append }) for a database/service adapter that validates the v1 schema and strips extra fields. Configure it as logging.store. Use decisionId as an idempotency key if your backend supports deduplication.

createJsonlDecisionStore(path) provides append and read(): Promise<GuardDecision[]>. It creates parent directories, appends UTF-8 JSON lines and flushes each append before resolving. A new instance can read previously stored records after a restart. Missing files read as an empty list. Invalid records cause read failures with a line number; incomplete final records stop both reading and appending. No repair or record deletion is automatic. Files use owner-only creation permissions where supported; filesystem access controls remain platform-dependent.

Calls targeting the same resolved path serialize within this process. Cross-process writers, symlink aliases and external rotation are not coordinated. Use one file per worker or a database adapter. The helper has no retention, encryption, rotation or deduplication. Reading loads the entire file. Manage file size and retention in the application. A crash or failed write may leave an incomplete record requiring operator recovery.

Delivery order: console event (if enabled), store append, then onDecision callback. Store and callback run independently of console enabled/level settings. Both are awaited and fail closed. Middleware delivery failures throw sanitized GuardOperationalError with code AUDIT_UNAVAILABLE, with no automatic retries or fallback to memory. Direct store calls retain their own errors. A callback failure does not undo an already persisted record. Pre-execution decisions are not receipts proving execution; tool response scans happen after the tool has already run.

The schema remains version 1. New reason codes are MODEL_FALLBACK_USED and MODEL_UNAVAILABLE; consumers matching reason-code enums should accept these additions. Built-in persisted events contain no prompt, arguments, classifier output or underlying error text. Policy versions and any custom events supplied by the application must also avoid sensitive values. Do not persist entire request metadata as an audit event.

Model selection and failure behavior

Setting Behavior
models.extractorFallback Optional fallback feature-extraction model
pii.fallback.model Explicit fallback token-classification model
pii.fallback.mode Defaults to primary PII mode
pii.fallback.labelMappings Independent mapping; primary mappings are not inherited

The primary is always attempted first. If loading or inference throws, try the configured fallback once. Non-array PII model output is treated as an invalid model result. A policy rejection, a low intent score, an empty PII result, a label mapping error, an audit-store failure or a downstream model/tool error does not cause a model fallback.

Startup uses the same configured primary/fallback choices for loading. A fallback is loaded lazily only if needed; successfully loading a primary does not guarantee future inference will succeed. Each subsequent operation tries the primary again. There is no timeout, retry loop, circuit breaker or process-global preferred fallback.

After recovery, MODEL_FALLBACK_USED is emitted for guard intent or pii with action allow. Here allow means the recovered model result is available; normal policy checks still apply and may block the request. If both attempts fail, MODEL_UNAVAILABLE is emitted with action block, then GuardModelError is thrown. Its code is MODEL_UNAVAILABLE; raw model exceptions are not attached. A failed audit delivery instead throws GuardOperationalError with code AUDIT_UNAVAILABLE. With no fallback configured, a model failure also emits MODEL_UNAVAILABLE and throws sanitized GuardModelError; the original exception and cause are not exposed.

Regex still runs alongside PII models, but if both models fail it is not used as a bypass. Tool-input scan failure prevents execution; tool-output scan failure cannot undo an already executed tool. Different models have different coverage and score distributions; evaluate recall/precision and thresholds before enabling a fallback. Classifier models require compatible q4 weights. The bundled download script prepares MiniLM and privacy-filter only, not configured backups.

Validation

Run npm run test:release2 for deterministic loading/inference fallback, audit failure, JSONL persistence/concurrency and Genkit integration tests. Real model accuracy, provider calls and Redis behavior are separate integration checks.

Model/tool request metadata adds piiEffectiveModel, piiEffectiveMode and piiUsedFallback to identify the detector actually used. Existing piiModel/piiMode metadata on model requests retain the configured primary values. Empty tool text skips detection and has no effective model. These metadata fields do not make the full request metadata safe to persist.

See 12.-Security-and-Operational-Errors for the complete error contract and retry boundaries.

Clone this wiki locally