-
Notifications
You must be signed in to change notification settings - Fork 0
Smart Routing
Smart Model Routing automatically selects the best model within a model group based on request complexity, cost constraints, and quality feedback. Instead of always routing to the highest-priority provider, the system classifies each request's complexity and dispatches it to the most cost-effective tier that can handle it well.
Smart routing is opt-in (enabled: false by default) and operates alongside the existing priority-based failover. When disabled, routing falls back to standard priority ordering.
Request ──► Context Capacity Filter ──► Complexity Classifier ──► Tier Selection ──► Model Dispatch
│ │ │
Reject if no model ┌─────┼─────┐ ┌──────┼──────┐
can fit the context │ │ │ │ │ │
Heuristic ML LLM Fast Balanced Powerful
│ │ │ │ │ │
└─────┼─────┘ └──────┼──────┘
│ │
▼ ▼
Complexity Score Best-fit Model
(0.0–1.0) (cost + quality)
-
Context filter — Exclude models whose
context_windowcannot safely fit the request (accounting for reserved output tokens, provider overhead, and safety margin) - Classify — Score the request for complexity (0.0 = trivial, 1.0 = highly complex)
- Map to tier — The score maps to a capability tier using configurable boundaries
- Budget check — Verify spend hasn't exceeded per-group limits; downgrade tier if needed
- Select model — Within the tier, pick the most cost-effective model with available capacity
- Cascade (optional) — If the response quality is insufficient, escalate to a higher tier
Each model in a model group can be assigned a capability tier:
| Tier | Use Case | Example Models |
|---|---|---|
| Fast | Simple queries, lookups, classification | GPT-3.5, Llama 3 8B, Gemini Flash |
| Balanced | Multi-step reasoning, code generation | GPT-4o, Claude Sonnet, Llama 3 70B |
| Powerful | Complex analysis, research, long-context | GPT-4.1, Claude Opus, o3 |
model_groups:
- name: "smart-group"
models:
- provider: "openai"
model: "gpt-4.1"
tier: powerful
context_window: 128000
specializations: [code_generation, factual_qa]
priority: 1
- provider: "openai"
model: "gpt-4o"
tier: balanced
context_window: 128000
priority: 2
- provider: "groq"
model: "llama3-8b-8192"
tier: fast
context_window: 8192
priority: 3The complexity score (0.0–1.0) maps to tiers via configurable boundaries:
smart_routing:
tier_boundaries:
fast_max: 0.33 # Scores 0.0–0.33 → Fast tier
balanced_max: 0.66 # Scores 0.33–0.66 → Balanced tier
# Scores 0.66–1.0 → Powerful tierSmart Routing has its own dedicated tab in the Admin Panel:

The tab is organized into sections:
| Section | Controls |
|---|---|
| Core Routing Policy | Enable/disable toggle, classifier mode dropdown (Heuristic / LLM / ONNX ML / Composite / Jev (System One) / Laya (System One)), cost/quality threshold slider, ONNX artifact directory, LLM classifier model |
| Jev Classifier / Laya Classifier | One fieldset per System One family: base URL, model, API key / env name, timeout, confidence gates, low-confidence policy, character budget, discovery TTL, retries, backoff, and the six dimension weights (see System One Classifiers) |
| Tier Boundaries and Context Safety | Fast/balanced boundary values, reserved output tokens, provider overhead tokens, context safety margin, unknown-context-window policy |
| Cascade and Streaming | Cascade enable toggle, max escalations, streaming cascade mode (Buffer / Early Signal), early signal tokens |
| Quality, Semantic Cache, and Optimizer | Quality evaluator toggle + threshold, semantic response cache toggle + similarity threshold, online optimizer toggle + EMA alpha |
| Budget Limits | JSON editor for per-model-group hourly/daily/monthly USD limits |
| ONNX Model Assets | Status indicator, "Check Assets" button, "Install Approved ONNX Assets" button |
| Safe Routing Simulation | Test routing decisions without sending a request — enter a model group and prompt, click "Simulate Without Generation" to see the tier/model selection |
Note: ML and Composite classifier modes require the
ml-routerbuild feature and a valid ONNX artifact installed in the configured directory.
The gateway offers six complexity classification strategies:
| Mode | Description | Requirements |
|---|---|---|
heuristic (default) |
Weighted signal analysis of request structure | None |
ml (ONNX ML) |
ONNX machine-learning model inference |
ml-router build feature + trained model |
llm |
Uses a configured LLM to assess complexity | Provider + model name |
composite |
Weighted blend of heuristic + ML scores |
ml-router build feature + trained model |
jev |
TypeSafe AI "System One" model returning typed scores + confidence across six complexity dimensions | Jev API key + endpoint |
laya |
Convai Innovations "System One" model; same typed scoring as jev over the shared wire protocol |
Laya API key + endpoint |
System One classifiers (
jev,laya): These modes call an external decision model that returns typed Score/Choice answers with confidence, instead of a parsed text label. They gate on confidence and fall back to the heuristic/ML/LLM chain when uncertain or unavailable. Both families share one implementation and differ only in defaults and confidence field. See System One Classifiers for setup, confidence gating, dimension weights, resilience, and metrics.
Build feature: The ML and Composite modes require the gateway to be compiled with
--features ml-router. The standard release binary includes this feature. The ONNX artifact directory must containmanifest.json,model.onnx,tokenizer.json, and an.onnxruntime/subdirectory.
Analyzes request signals with configurable weights (must sum to 1.0):
smart_routing:
classifier: heuristic
heuristic_weights:
message_count: 0.167 # Number of conversation messages
token_estimate: 0.167 # Estimated input token count
code_blocks: 0.167 # Presence of fenced code blocks
tool_calls: 0.167 # Number of tool definitions
math_expressions: 0.167 # Mathematical notation detected
reasoning_keywords: 0.167 # Keywords indicating complex reasoningUses a pre-trained ONNX model for complexity scoring. The artifact directory must contain the model, tokenizer, manifest, and runtime:
smart_routing:
classifier: ml
ml_model_path: ./models/smart-routing # Directory containing manifest.json, model.onnx, etc.Install the approved ONNX assets from the Admin Panel's Smart Routing tab → ONNX Model Assets → "Install Approved ONNX Assets", or provide your own trained artifacts.
Delegates complexity assessment to a configured model:
smart_routing:
classifier: llm
classifier_model: gpt-4o-miniBlends heuristic and ML scores:
smart_routing:
classifier: composite
ml_model_path: ./models/smart-routing # ONNX artifact directory
composite_weights:
heuristic: 0.4
ml: 0.6When enabled, cascade monitors the response quality from a lower tier and escalates to a higher tier if the response is insufficient:
smart_routing:
cascade:
enabled: true
max_escalations: 2 # Maximum tier jumps (1–2)
min_response_tokens: 50 # Minimum tokens before evaluating quality
early_signal_tokens: 100 # Tokens to inspect for early escalation signals| Mode | Behavior |
|---|---|
buffer (default) |
Buffer the response, evaluate quality, re-dispatch if needed |
early_signal |
Inspect the first N tokens during streaming; escalate before completion |
smart_routing:
streaming_cascade_mode: buffer # buffer | early_signalThe online optimizer continuously adjusts tier boundaries based on observed quality scores:
smart_routing:
online_optimizer:
enabled: true
alpha: 0.1 # EMA learning rate (0.0–1.0)
interval_secs: 3600 # Optimization interval (1–604800)
state_path: ./smart_routing_state.json
quality_threshold: 0.7 # Minimum acceptable quality (0.0–1.0)When quality drops below the threshold for a tier, the optimizer shifts boundaries to route more requests to higher tiers. When quality exceeds the threshold consistently, boundaries relax to save cost.
Evaluates response quality to feed the cascade and optimizer:
smart_routing:
quality_evaluator:
enabled: true
threshold: 0.7 # Quality score below this triggers escalationSmart routing includes its own semantic cache (separate from the gateway's response cache) that caches complexity classification decisions:
smart_routing:
semantic_cache:
enabled: true
similarity_threshold: 0.9
max_entries: 10000
ttl_secs: 3600
embedding_model: builtin-minilm
min_quality_score: 0.7When a new request is semantically similar to a previously classified request, the cached tier decision is reused — skipping classification entirely.
Per-model-group spend limits for smart routing decisions:
smart_routing:
budget_limits:
smart-group:
hourly_limit_usd: 5.0
daily_limit_usd: 50.0
monthly_limit_usd: 500.0When a budget limit is reached, the router favors cheaper tiers or falls back to standard priority routing.
Run controlled experiments comparing two routing policies:
smart_routing:
ab_test:
variant_percentage: 0.2 # 20% of traffic uses the variant policy
control:
classifier: heuristic
cost_quality_threshold: 0.5
variant:
classifier: composite
ml_model_path: ./models/new_classifier.onnx
cost_quality_threshold: 0.7The variant_percentage (0.0–1.0) controls what fraction of requests use the variant policy. Quality metrics are tracked separately for each arm.
Configure offline training jobs for the ML classifier:
smart_routing:
training:
dataset_path: ./training_data.jsonl
learning_rate: 0.001
batch_size: 32
epochs: 10
augmentation: trueTraining data is collected from the online optimizer's quality observations. The trained model replaces the existing ml_model_path after validation.
Each model group can override any global smart routing setting:
smart_routing:
enabled: true
classifier: heuristic
model_group_overrides:
coding-group:
classifier: composite
ml_model_path: ./models/code_classifier.onnx
cost_quality_threshold: 0.8
cascade:
enabled: true
max_escalations: 1
simple-group:
enabled: false # Disable smart routing for this groupResolution: per-group override fields replace the corresponding global fields; omitted fields inherit from the global config.
Smart routing considers each model's declared context_window when selecting targets. Before any tier selection occurs, models are filtered by token capacity:
-
Estimate total requirement:
input_tokens + reserved_output_tokens + provider_overhead_tokens + context_safety_margin_tokens -
Filter candidates: Models whose
context_windowis less than the requirement are excluded -
Unknown capacity policy: Models with
context_window: 0(unknown) are excluded unlessallow_unknown_context_window: true - HTTP 413 on exhaustion: If no candidate survives filtering, the gateway returns HTTP 413 with the estimated requirement and largest known context
smart_routing:
reserved_output_tokens: 1024 # Tokens reserved for model output (default: 1024)
provider_overhead_tokens: 64 # Tokens for provider-injected content (default: 64)
context_safety_margin_tokens: 256 # Additional headroom (default: 256)
allow_unknown_context_window: false # Reject models with context_window: 0This ensures the router never sends a request to a model that can't fit it, avoiding wasteful context-overflow errors and failover delays.
smart_routing:
enabled: false
classifier: heuristic # heuristic | ml | llm | composite | jev | laya
ml_model_path: null # Path to ONNX artifact directory
classifier_model: null # LLM model name (for llm classifier)
# jev: / laya: # System One blocks, see System One Classifiers page
cost_quality_threshold: 0.5 # Cost vs quality tradeoff (0.0–1.0; 0=cost, 1=quality)
tier_boundaries:
fast_max: 0.33
balanced_max: 0.66
reserved_output_tokens: 1024 # Tokens reserved for output generation
provider_overhead_tokens: 64 # Tokens for provider-injected content
context_safety_margin_tokens: 256 # Additional safety headroom
allow_unknown_context_window: false
cascade:
enabled: false
max_escalations: 2
min_response_tokens: 50
early_signal_tokens: 100
streaming_cascade_mode: buffer # buffer | early_signal
heuristic_weights: # Must sum to 1.0
message_count: 0.167
token_estimate: 0.167
code_blocks: 0.167
tool_calls: 0.167
math_expressions: 0.167
reasoning_keywords: 0.167
composite_weights: # Must sum to 1.0
heuristic: 0.4
ml: 0.6
online_optimizer:
enabled: false
alpha: 0.01
interval_secs: 3600
state_path: null
quality_threshold: 0.7
quality_evaluator:
enabled: false
threshold: 0.3
semantic_cache:
enabled: false
similarity_threshold: 0.95
max_entries: 10000
ttl_secs: 3600
embedding_model: builtin-minilm
min_quality_score: 0.7
budget_limits: {} # Model group name → {hourly_limit_usd, daily_limit_usd, monthly_limit_usd}
ab_test: null
training:
dataset_path: null
learning_rate: 0.001
batch_size: 32
epochs: 10
augmentation: false
model_group_overrides: {}Models can declare task categories they excel at:
models:
- provider: openai
model: gpt-4.1
tier: powerful
specializations: [code_generation, factual_qa]Available task types allow the router to prefer models with matching specializations when the request is classified into a particular task category.
The Admin Panel includes a Safe Routing Simulation tool that lets you test routing decisions without sending a request to any provider:
- Enter a Model Group name (must match a configured group with tiered models)
- Enter a Prompt (the user message content to classify)
- Click "Simulate Without Generation"
The simulator runs the full classification pipeline (heuristic scoring, tier selection, context filtering, budget check) and returns:
- Complexity score and task type
- Selected tier and model
- Candidates excluded for context capacity
- Whether cascade would trigger
This is useful for:
- Validating tier boundaries before enabling smart routing in production
- Testing custom heuristic weights against real prompts
- Verifying that expensive models are only selected for genuinely complex requests
The Smart Routing tab includes an ONNX Model Assets section for managing the ML classifier model:
| Action | Purpose |
|---|---|
| Check Assets | Validates the artifact directory: checks for manifest.json, model.onnx, tokenizer.json, and .onnxruntime/
|
| Install Approved ONNX Assets | Downloads and installs the pinned, approved classifier model and runtime into the configured directory |
The assets are installed server-side (not in the browser). Docker users should mount a writable persistent volume at the artifact directory path.
- System One Classifiers (Jev & Laya) — calibrated decision-model classifiers with typed scores and confidence gating
- Routing & Failover — standard priority-based routing (used as fallback)
- Configuration — full config reference
- Admin Panel & Dashboard — monitoring