-
Notifications
You must be signed in to change notification settings - Fork 0
Smart Routing
Smart Model Routing automatically selects the best model within a model group based on request complexity, cost constraints, and quality feedback. Instead of always routing to the highest-priority provider, the system classifies each request's complexity and dispatches it to the most cost-effective tier that can handle it well.
Smart routing is opt-in (enabled: false by default) and operates alongside the existing priority-based failover. When disabled, routing falls back to standard priority ordering.
Request ──► Complexity Classifier ──► Tier Selection ──► Model Dispatch
│ │
┌───────┼───────┐ ┌───────┼───────┐
│ │ │ │ │ │
Heuristic ML LLM Fast Balanced Powerful
│ │ │ │ │ │
└───────┼───────┘ └───────┼───────┘
│ │
▼ ▼
Complexity Score Best-fit Model
(0.0–1.0) (cost + quality)
- Classify — The request is scored for complexity (0.0 = trivial, 1.0 = highly complex)
- Map to tier — The score is mapped to a capability tier using configurable boundaries
- Select model — Within the tier, the most cost-effective model with available capacity is chosen
- Cascade (optional) — If the response quality is insufficient, escalate to a higher tier
Each model in a model group can be assigned a capability tier:
| Tier | Use Case | Example Models |
|---|---|---|
| Fast | Simple queries, lookups, classification | GPT-3.5, Llama 3 8B, Gemini Flash |
| Balanced | Multi-step reasoning, code generation | GPT-4o, Claude Sonnet, Llama 3 70B |
| Powerful | Complex analysis, research, long-context | GPT-4.1, Claude Opus, o3 |
model_groups:
- name: "smart-group"
models:
- provider: "openai"
model: "gpt-4.1"
tier: powerful
context_window: 128000
specializations: [code_generation, factual_qa]
priority: 1
- provider: "openai"
model: "gpt-4o"
tier: balanced
context_window: 128000
priority: 2
- provider: "groq"
model: "llama3-8b-8192"
tier: fast
context_window: 8192
priority: 3The complexity score (0.0–1.0) maps to tiers via configurable boundaries:
smart_routing:
tier_boundaries:
fast_max: 0.3 # Scores 0.0–0.3 → Fast tier
balanced_max: 0.7 # Scores 0.3–0.7 → Balanced tier
# Scores 0.7–1.0 → Powerful tierThe gateway offers four complexity classification strategies:
| Mode | Description | Requirements |
|---|---|---|
heuristic (default) |
Weighted signal analysis of request structure | None |
ml |
ONNX machine-learning model inference | Trained model file |
llm |
Uses a configured LLM to assess complexity | Provider + model |
composite |
Weighted blend of heuristic + ML scores | Heuristic + ML model |
Analyzes request signals with configurable weights (must sum to 1.0):
smart_routing:
classifier: heuristic
heuristic_weights:
message_count: 0.167 # Number of conversation messages
token_estimate: 0.167 # Estimated input token count
code_blocks: 0.167 # Presence of fenced code blocks
tool_calls: 0.167 # Number of tool definitions
math_expressions: 0.167 # Mathematical notation detected
reasoning_keywords: 0.167 # Keywords indicating complex reasoningUses a pre-trained ONNX model for complexity scoring:
smart_routing:
classifier: ml
ml_model_path: ./models/complexity_classifier.onnxDelegates complexity assessment to a configured model:
smart_routing:
classifier: llm
classifier_model: gpt-4o-miniBlends heuristic and ML scores:
smart_routing:
classifier: composite
ml_model_path: ./models/complexity_classifier.onnx
composite_weights:
heuristic: 0.4
ml: 0.6When enabled, cascade monitors the response quality from a lower tier and escalates to a higher tier if the response is insufficient:
smart_routing:
cascade:
enabled: true
max_escalations: 2 # Maximum tier jumps (1–2)
min_response_tokens: 50 # Minimum tokens before evaluating quality
early_signal_tokens: 100 # Tokens to inspect for early escalation signals| Mode | Behavior |
|---|---|
buffer (default) |
Buffer the response, evaluate quality, re-dispatch if needed |
early_signal |
Inspect the first N tokens during streaming; escalate before completion |
smart_routing:
streaming_cascade_mode: buffer # buffer | early_signalThe online optimizer continuously adjusts tier boundaries based on observed quality scores:
smart_routing:
online_optimizer:
enabled: true
alpha: 0.1 # EMA learning rate (0.0–1.0)
interval_secs: 3600 # Optimization interval (1–604800)
state_path: ./smart_routing_state.json
quality_threshold: 0.7 # Minimum acceptable quality (0.0–1.0)When quality drops below the threshold for a tier, the optimizer shifts boundaries to route more requests to higher tiers. When quality exceeds the threshold consistently, boundaries relax to save cost.
Evaluates response quality to feed the cascade and optimizer:
smart_routing:
quality_evaluator:
enabled: true
threshold: 0.7 # Quality score below this triggers escalationSmart routing includes its own semantic cache (separate from the gateway's response cache) that caches complexity classification decisions:
smart_routing:
semantic_cache:
enabled: true
similarity_threshold: 0.9
max_entries: 10000
ttl_secs: 3600
embedding_model: builtin-minilm
min_quality_score: 0.7When a new request is semantically similar to a previously classified request, the cached tier decision is reused — skipping classification entirely.
Per-model-group spend limits for smart routing decisions:
smart_routing:
budget_limits:
smart-group:
hourly_limit_usd: 5.0
daily_limit_usd: 50.0
monthly_limit_usd: 500.0When a budget limit is reached, the router favors cheaper tiers or falls back to standard priority routing.
Run controlled experiments comparing two routing policies:
smart_routing:
ab_test:
variant_percentage: 0.2 # 20% of traffic uses the variant policy
control:
classifier: heuristic
cost_quality_threshold: 0.5
variant:
classifier: composite
ml_model_path: ./models/new_classifier.onnx
cost_quality_threshold: 0.7The variant_percentage (0.0–1.0) controls what fraction of requests use the variant policy. Quality metrics are tracked separately for each arm.
Configure offline training jobs for the ML classifier:
smart_routing:
training:
dataset_path: ./training_data.jsonl
learning_rate: 0.001
batch_size: 32
epochs: 10
augmentation: trueTraining data is collected from the online optimizer's quality observations. The trained model replaces the existing ml_model_path after validation.
Each model group can override any global smart routing setting:
smart_routing:
enabled: true
classifier: heuristic
model_group_overrides:
coding-group:
classifier: composite
ml_model_path: ./models/code_classifier.onnx
cost_quality_threshold: 0.8
cascade:
enabled: true
max_escalations: 1
simple-group:
enabled: false # Disable smart routing for this groupResolution: per-group override fields replace the corresponding global fields; omitted fields inherit from the global config.
Smart routing considers each model's declared context_window when selecting targets:
- Requests exceeding a model's context window skip that model
-
reserved_output_tokens(default: 4096) is subtracted from the available window -
provider_overhead_tokens(default: 256) accounts for provider-injected content -
context_safety_margin_tokens(default: 512) provides additional headroom
smart_routing:
reserved_output_tokens: 4096
provider_overhead_tokens: 256
context_safety_margin_tokens: 512
allow_unknown_context_window: false # Reject models with context_window: 0smart_routing:
enabled: false
classifier: heuristic # heuristic | ml | llm | composite
ml_model_path: null # Path to ONNX classifier model
classifier_model: null # LLM model name (for llm classifier)
cost_quality_threshold: 0.5 # Cost vs quality tradeoff (0.0–1.0)
tier_boundaries:
fast_max: 0.3
balanced_max: 0.7
cascade:
enabled: false
max_escalations: 2
min_response_tokens: 50
early_signal_tokens: 100
streaming_cascade_mode: buffer # buffer | early_signal
heuristic_weights: # Must sum to 1.0
message_count: 0.167
token_estimate: 0.167
code_blocks: 0.167
tool_calls: 0.167
math_expressions: 0.167
reasoning_keywords: 0.167
composite_weights: # Must sum to 1.0
heuristic: 0.4
ml: 0.6
online_optimizer:
enabled: false
alpha: 0.1
interval_secs: 3600
state_path: null
quality_threshold: 0.7
quality_evaluator:
enabled: false
threshold: 0.7
semantic_cache:
enabled: false
similarity_threshold: 0.9
max_entries: 10000
ttl_secs: 3600
embedding_model: builtin-minilm
min_quality_score: 0.7
budget_limits: {}
ab_test: null
training:
dataset_path: null
learning_rate: 0.001
batch_size: 32
epochs: 10
augmentation: false
reserved_output_tokens: 4096
provider_overhead_tokens: 256
context_safety_margin_tokens: 512
allow_unknown_context_window: false
model_group_overrides: {}Models can declare task categories they excel at:
models:
- provider: openai
model: gpt-4.1
tier: powerful
specializations: [code_generation, factual_qa]Available task types allow the router to prefer models with matching specializations when the request is classified into a particular task category.
- Routing & Failover — standard priority-based routing (used as fallback)
- Configuration — full config reference
- Admin Panel & Dashboard — monitoring