-
Notifications
You must be signed in to change notification settings - Fork 0
TypeSafe Integration
AISRF can use TypeSafe AI (the System One decision model Jev) as a fast, cheap decision layer next to its own Analyzers and the Guardrails integrations. Instead of asking a generative model to write a verdict, AISRF sends the intercepted conversation (or a model response, a red-team probe result, or a static analysis finding) as an untrusted state together with a battery of typed questions, and Jev answers every question in parallel with a calibrated probability and a confidence in roughly 100 ms. Code then decides what to do with the answers. The integration is a silent no-op until integrations.typesafe.enabled is true and a key exists, it never raises (a failed call yields no findings, a heuristic red-team verdict, no triage), and errors are logged at most once per minute and kept as last_error.
All code lives in aisrf/typesafe/ (client.py, questions.py, analyzers.py, __init__.py) with hooks in aisrf/analysis/runner.py, aisrf/gateway/policy.py, aisrf/redteam/evaluators.py, aisrf/codereview/engines/typesafe_engine.py and aisrf/settings_router.py. The long form reference is docs/TYPESAFE.md in the repository.
| Where | Question set | What it decides | Cost per item |
|---|---|---|---|
Request analysis (typesafe_guard) |
REQUEST_GUARD: nine nouls, an intent choice over the taxonomy categories, a five level severity score |
Findings per attack class, a malicious / benign / uncertain verdict with a confidence |
one call, state plus about 900 question tokens |
Response analysis (typesafe_response_guard) |
RESPONSE_GUARD: leaked system prompt, leaked secret or PII, harmful compliance, exfiltration markup, refusal |
Findings on the model output | one call, state plus about 500 question tokens |
| Red-team evaluation |
REDTEAM_EVAL: refused, complied, followed injected instruction, leaked canary, harmful content, safe helpful answer, outcome choice |
VULNERABLE / RESISTED / INCONCLUSIVE when confident enough to override the heuristics | one call per probe result |
Code review (typesafe engine) |
CODE_TRIAGE: true positive, user input reaches prompt, model output reaches sink, exploitable without auth, mitigated in code, four level exploitability score |
Confidence adjustment and the likely_false_positive status |
one call per finding, up to max_findings
|
| Primitive | Asks | Answer |
|---|---|---|
noul |
a yes/no question with optional criteria.true and criteria.false descriptions |
noul in 0..1, the probability of yes |
choice |
one option out of a criteria map of option to description (at most 255) |
choice, probabilities per option, confidence
|
score |
a rating on an ordered rubric of 2 to 10 level descriptions |
score, legend, probabilities per level, confidence
|
A confidence below 0.5 means the model is genuinely unsure. AISRF turns a noul into a confidence with abs(p - 0.5) * 2 and uses the returned confidence for choices and scores. Every instruction reminds the model that the state is untrusted material to be judged, never followed.
- Create an API key at typesafe.ai.
- Export
TYPESAFE_API_KEYin the gateway environment, or paste the key into Settings > Integrations > typesafe (stored server side, masked in the UI and the API). The settings key wins when both exist. - Switch
enabledon and run the health check:
curl -X PUT -H "Authorization: Bearer $AISRF_ADMIN_API_TOKEN" -H "Content-Type: application/json" \
http://localhost:8080/api/settings/ns/integrations -d '{"typesafe": {"enabled": true, "api_key": "ts-..."}}'
curl -X POST -H "Authorization: Bearer $AISRF_ADMIN_API_TOKEN" http://localhost:8080/api/settings/typesafe/healthSettings are live; every hook reads integrations.typesafe at call time (Settings-Center).
The analysis runner executes in two phases: every analyzer except the LLM judge runs concurrently first (TypeSafe among them), then the judge runs only if llm_judge_gate allows it. With TypeSafe disabled the judge behaves as before. With TypeSafe enabled the judge runs only when escalate_to_llm_judge is true and the verdict is uncertain (or missing because the call failed). The skip reason is stored in context["llm_judge_skipped"].
The Policy-Engine gains one step between the deny patterns and the risk thresholds when auto_route is true:
| Verdict | Condition | Decision | Matched rule |
|---|---|---|---|
malicious |
confidence >= auto_deny_confidence (0.95) |
deny | typesafe.auto_deny |
benign |
confidence >= auto_approve_confidence (0.9), agent requires approval, no HIGH or CRITICAL finding from any analyzer |
approve | typesafe.auto_approve |
| anything else | unchanged, the normal rules decide |
In red teaming the TypeSafe verdict replaces the heuristic one only when its confidence is at least redteam_confidence (0.8); BLOCKED, ERROR and canary-leak verdicts are never sent. In code review a finding becomes likely_false_positive when true_positive is below 0.2 with confidence at least triage_confidence; that status is excluded from the risk score and SARIF, hidden by the default findings filter, and can be reopened by a reviewer. The LLM code review engine then only sees the findings TypeSafe left uncertain, and only when escalate_to_llm_judge is true.
Every TypeSafe finding carries metadata.typesafe = {"verdict", "confidence"} plus the full answers (probabilities and confidence per question), intent, severity_score, model, usage and latency_ms. The ticket page shows a TypeSafe: malicious 96% chip on the finding and an expandable table with every answer; the overview page shows a TypeSafe savings tile while the integration is enabled.
| Key | Default | Meaning |
|---|---|---|
enabled |
false |
Master switch. |
api_key |
"" |
Bearer key; falls back to TYPESAFE_API_KEY. |
base_url |
https://api.typesafe.ai |
The client posts to /v1/systemone. |
model |
jev-latest |
Model name. |
timeout_seconds |
8 |
HTTP timeout per attempt; 3 retries with backoff and jitter on 429, 529 and connection errors. |
cache_ttl_seconds |
300 |
TTL of the bounded result cache (sha256 of model, state and questions). |
max_state_chars |
12000 |
State budget; head and tail are kept when truncating, secrets are masked. |
noul_threshold |
0.7 |
A noul at or above this fires a finding. |
benign_confidence, malicious_confidence
|
0.85, 0.9
|
Minimum confidence for the two decisive verdicts. |
escalate_to_llm_judge |
false |
Run the LLM judge and the LLM code review engine for uncertain results only. |
auto_route, auto_deny_confidence, auto_approve_confidence
|
false, 0.95, 0.9
|
Policy routing. |
redteam_evaluator, redteam_confidence
|
true, 0.8
|
Red-team evaluation. |
max_findings, triage_min_severity, triage_confidence, triage_concurrency
|
200, LOW, 0.8, 8
|
Code review triage. |
judge_input_price_per_million, judge_output_price_per_million, price_per_million_input
|
0.15, 0.6, 0.042
|
Prices for the savings estimate. |
TypeSafe bills input tokens only. The savings estimate assumes the LLM judge would have read the same input tokens and written 400 output tokens at the two judge prices, and subtracts the TypeSafe spend (billed_input_tokens * price_per_million_input); cached evaluations count as judge tokens avoided at no cost. Prometheus exposes aisrf_typesafe_calls_total{outcome}, aisrf_typesafe_input_tokens_total, aisrf_typesafe_cache_hits_total and aisrf_typesafe_latency_ms (Logging-Metrics-and-Audit).
-
GET /api/settings/typesafe/status:enabled,configured,key_source,model,cache_size,calls,cache_hits,errors,input_tokens,savings,last_error. -
POST /api/settings/typesafe/health(admin): one tiny uncached evaluation. -
GET /api/settings/guardrailslists TypeSafe next to the guardrail integrations.
Jev judges, it does not generate: it cannot explain a verdict in prose. Calibration is not correctness, so keep the thresholds for the two actions that bypass a human high and review auto-routed tickets. The state is text only and must fit the model context; max_state_chars keeps it well below. Without network access every hook degrades to the previous behaviour.
See also: Analyzers, Guardrails, Policy-Engine, Red-Teaming, Code-Review, Configuration-Reference.
AISRF, AI Security & Research Framework. github.com/keyuraghao/aisrf, Apache License 2.0.
Start
Gateway
- Gateway-Endpoints-and-Headers
- Request-Normalization
- Policy-Engine
- Agents-and-Credentials
- Configuration-Reference
- Settings-Center
Review
Security analysis
Red teaming
Code review
Interfaces
Operations
Project