Skip to content

Smart Routing

fdanobey edited this page Oct 3, 2026 · 4 revisions

Smart Model Routing

Smart Model Routing automatically selects the best model within a model group based on request complexity, cost constraints, and quality feedback. Instead of always routing to the highest-priority provider, the system classifies each request's complexity and dispatches it to the most cost-effective tier that can handle it well.

Smart routing is opt-in (enabled: false by default) and operates alongside the existing priority-based failover. When disabled, routing falls back to standard priority ordering.


How It Works

Request ──► Context Capacity Filter ──► Complexity Classifier ──► Tier Selection ──► Model Dispatch
                    │                           │                       │
            Reject if no model            ┌─────┼─────┐         ┌──────┼──────┐
            can fit the context           │     │     │         │      │      │
                                       Heuristic ML  LLM     Fast  Balanced Powerful
                                          │     │     │         │      │      │
                                          └─────┼─────┘         └──────┼──────┘
                                                │                      │
                                                ▼                      ▼
                                        Complexity Score         Best-fit Model
                                          (0.0–1.0)           (cost + quality)
  1. Context filter — Exclude models whose context_window cannot safely fit the request (accounting for reserved output tokens, provider overhead, and safety margin)
  2. Classify — Score the request for complexity (0.0 = trivial, 1.0 = highly complex)
  3. Map to tier — The score maps to a capability tier using configurable boundaries
  4. Budget check — Verify spend hasn't exceeded per-group limits; downgrade tier if needed
  5. Select model — Within the tier, pick the most cost-effective model with available capacity
  6. Cascade (optional) — If the response quality is insufficient, escalate to a higher tier

Model Tiers

Each model in a model group can be assigned a capability tier:

Tier Use Case Example Models
Fast Simple queries, lookups, classification GPT-3.5, Llama 3 8B, Gemini Flash
Balanced Multi-step reasoning, code generation GPT-4o, Claude Sonnet, Llama 3 70B
Powerful Complex analysis, research, long-context GPT-4.1, Claude Opus, o3
model_groups:
  - name: "smart-group"
    models:
      - provider: "openai"
        model: "gpt-4.1"
        tier: powerful
        context_window: 128000
        specializations: [code_generation, factual_qa]
        priority: 1

      - provider: "openai"
        model: "gpt-4o"
        tier: balanced
        context_window: 128000
        priority: 2

      - provider: "groq"
        model: "llama3-8b-8192"
        tier: fast
        context_window: 8192
        priority: 3

Tier Boundaries

The complexity score (0.0–1.0) maps to tiers via configurable boundaries:

smart_routing:
  tier_boundaries:
    fast_max: 0.33       # Scores 0.0–0.33 → Fast tier
    balanced_max: 0.66   # Scores 0.33–0.66 → Balanced tier
                         # Scores 0.66–1.0 → Powerful tier

Admin Panel

Smart Routing has its own dedicated tab in the Admin Panel:

Smart Routing Admin

The tab is organized into sections:

Section Controls
Core Routing Policy Enable/disable toggle, classifier mode dropdown (Heuristic / LLM / ONNX ML / Composite / Jev (System One) / Laya (System One)), cost/quality threshold slider, ONNX artifact directory, LLM classifier model
Jev Classifier / Laya Classifier One fieldset per System One family: base URL, model, API key / env name, timeout, confidence gates, low-confidence policy, character budget, discovery TTL, retries, backoff, and the six dimension weights (see System One Classifiers)
Tier Boundaries and Context Safety Fast/balanced boundary values, reserved output tokens, provider overhead tokens, context safety margin, unknown-context-window policy
Cascade and Streaming Cascade enable toggle, max escalations, streaming cascade mode (Buffer / Early Signal), early signal tokens
Quality, Semantic Cache, and Optimizer Quality evaluator toggle + threshold, semantic response cache toggle + similarity threshold, online optimizer toggle + EMA alpha
Budget Limits JSON editor for per-model-group hourly/daily/monthly USD limits
ONNX Model Assets Status indicator, "Check Assets" button, "Install Approved ONNX Assets" button
Safe Routing Simulation Test routing decisions without sending a request — enter a model group and prompt, click "Simulate Without Generation" to see the tier/model selection

Note: ML and Composite classifier modes require the ml-router build feature and a valid ONNX artifact installed in the configured directory.


Classifier Modes

The gateway offers six complexity classification strategies:

Mode Description Requirements
heuristic (default) Weighted signal analysis of request structure None
ml (ONNX ML) ONNX machine-learning model inference ml-router build feature + trained model
llm Uses a configured LLM to assess complexity Provider + model name
composite Weighted blend of heuristic + ML scores ml-router build feature + trained model
jev TypeSafe AI "System One" model returning typed scores + confidence across six complexity dimensions Jev API key + endpoint
laya Convai Innovations "System One" model; same typed scoring as jev over the shared wire protocol Laya API key + endpoint

System One classifiers (jev, laya): These modes call an external decision model that returns typed Score/Choice answers with confidence, instead of a parsed text label. They gate on confidence and fall back to the heuristic/ML/LLM chain when uncertain or unavailable. Both families share one implementation and differ only in defaults and confidence field. See System One Classifiers for setup, confidence gating, dimension weights, resilience, and metrics.

Build feature: The ML and Composite modes require the gateway to be compiled with --features ml-router. The standard release binary includes this feature. The ONNX artifact directory must contain manifest.json, model.onnx, tokenizer.json, and an .onnxruntime/ subdirectory.

Heuristic Classifier

Analyzes request signals with configurable weights (must sum to 1.0):

smart_routing:
  classifier: heuristic
  heuristic_weights:
    message_count: 0.167       # Number of conversation messages
    token_estimate: 0.167      # Estimated input token count
    code_blocks: 0.167         # Presence of fenced code blocks
    tool_calls: 0.167          # Number of tool definitions
    math_expressions: 0.167    # Mathematical notation detected
    reasoning_keywords: 0.167  # Keywords indicating complex reasoning

ML Classifier (ONNX)

Uses a pre-trained ONNX model for complexity scoring. The artifact directory must contain the model, tokenizer, manifest, and runtime:

smart_routing:
  classifier: ml
  ml_model_path: ./models/smart-routing   # Directory containing manifest.json, model.onnx, etc.

Install the approved ONNX assets from the Admin Panel's Smart Routing tab → ONNX Model Assets → "Install Approved ONNX Assets", or provide your own trained artifacts.

LLM Classifier

Delegates complexity assessment to a configured model:

smart_routing:
  classifier: llm
  classifier_model: gpt-4o-mini

Composite Classifier

Blends heuristic and ML scores:

smart_routing:
  classifier: composite
  ml_model_path: ./models/smart-routing   # ONNX artifact directory
  composite_weights:
    heuristic: 0.4
    ml: 0.6

Cascade Escalation

When enabled, cascade monitors the response quality from a lower tier and escalates to a higher tier if the response is insufficient:

smart_routing:
  cascade:
    enabled: true
    max_escalations: 2          # Maximum tier jumps (1–2)
    min_response_tokens: 50     # Minimum tokens before evaluating quality
    early_signal_tokens: 100    # Tokens to inspect for early escalation signals

Streaming Cascade Mode

Mode Behavior
buffer (default) Buffer the response, evaluate quality, re-dispatch if needed
early_signal Inspect the first N tokens during streaming; escalate before completion
smart_routing:
  streaming_cascade_mode: buffer   # buffer | early_signal

Online Optimizer

The online optimizer continuously adjusts tier boundaries based on observed quality scores:

smart_routing:
  online_optimizer:
    enabled: true
    alpha: 0.1                  # EMA learning rate (0.0–1.0)
    interval_secs: 3600         # Optimization interval (1–604800)
    state_path: ./smart_routing_state.json
    quality_threshold: 0.7      # Minimum acceptable quality (0.0–1.0)

When quality drops below the threshold for a tier, the optimizer shifts boundaries to route more requests to higher tiers. When quality exceeds the threshold consistently, boundaries relax to save cost.


Quality Evaluator

Evaluates response quality to feed the cascade and optimizer:

smart_routing:
  quality_evaluator:
    enabled: true
    threshold: 0.7              # Quality score below this triggers escalation

Semantic Cache

Smart routing includes its own semantic cache (separate from the gateway's response cache) that caches complexity classification decisions:

smart_routing:
  semantic_cache:
    enabled: true
    similarity_threshold: 0.9
    max_entries: 10000
    ttl_secs: 3600
    embedding_model: builtin-minilm
    min_quality_score: 0.7

When a new request is semantically similar to a previously classified request, the cached tier decision is reused — skipping classification entirely.


Budget Limits

Per-model-group spend limits for smart routing decisions:

smart_routing:
  budget_limits:
    smart-group:
      hourly_limit_usd: 5.0
      daily_limit_usd: 50.0
      monthly_limit_usd: 500.0

When a budget limit is reached, the router favors cheaper tiers or falls back to standard priority routing.


A/B Testing

Run controlled experiments comparing two routing policies:

smart_routing:
  ab_test:
    variant_percentage: 0.2     # 20% of traffic uses the variant policy
    control:
      classifier: heuristic
      cost_quality_threshold: 0.5
    variant:
      classifier: composite
      ml_model_path: ./models/new_classifier.onnx
      cost_quality_threshold: 0.7

The variant_percentage (0.0–1.0) controls what fraction of requests use the variant policy. Quality metrics are tracked separately for each arm.


Training

Configure offline training jobs for the ML classifier:

smart_routing:
  training:
    dataset_path: ./training_data.jsonl
    learning_rate: 0.001
    batch_size: 32
    epochs: 10
    augmentation: true

Training data is collected from the online optimizer's quality observations. The trained model replaces the existing ml_model_path after validation.


Per-Model-Group Overrides

Each model group can override any global smart routing setting:

smart_routing:
  enabled: true
  classifier: heuristic

  model_group_overrides:
    coding-group:
      classifier: composite
      ml_model_path: ./models/code_classifier.onnx
      cost_quality_threshold: 0.8
      cascade:
        enabled: true
        max_escalations: 1
    simple-group:
      enabled: false    # Disable smart routing for this group

Resolution: per-group override fields replace the corresponding global fields; omitted fields inherit from the global config.


Context Window Awareness

Smart routing considers each model's declared context_window when selecting targets. Before any tier selection occurs, models are filtered by token capacity:

  1. Estimate total requirement: input_tokens + reserved_output_tokens + provider_overhead_tokens + context_safety_margin_tokens
  2. Filter candidates: Models whose context_window is less than the requirement are excluded
  3. Unknown capacity policy: Models with context_window: 0 (unknown) are excluded unless allow_unknown_context_window: true
  4. HTTP 413 on exhaustion: If no candidate survives filtering, the gateway returns HTTP 413 with the estimated requirement and largest known context
smart_routing:
  reserved_output_tokens: 1024        # Tokens reserved for model output (default: 1024)
  provider_overhead_tokens: 64        # Tokens for provider-injected content (default: 64)
  context_safety_margin_tokens: 256   # Additional headroom (default: 256)
  allow_unknown_context_window: false # Reject models with context_window: 0

This ensures the router never sends a request to a model that can't fit it, avoiding wasteful context-overflow errors and failover delays.


Configuration Reference

smart_routing:
  enabled: false
  classifier: heuristic              # heuristic | ml | llm | composite | jev | laya
  ml_model_path: null                # Path to ONNX artifact directory
  classifier_model: null             # LLM model name (for llm classifier)
  # jev: / laya:                     # System One blocks, see System One Classifiers page
  cost_quality_threshold: 0.5        # Cost vs quality tradeoff (0.0–1.0; 0=cost, 1=quality)

  tier_boundaries:
    fast_max: 0.33
    balanced_max: 0.66

  reserved_output_tokens: 1024       # Tokens reserved for output generation
  provider_overhead_tokens: 64       # Tokens for provider-injected content
  context_safety_margin_tokens: 256  # Additional safety headroom
  allow_unknown_context_window: false

  cascade:
    enabled: false
    max_escalations: 2
    min_response_tokens: 50
    early_signal_tokens: 100

  streaming_cascade_mode: buffer     # buffer | early_signal

  heuristic_weights:                 # Must sum to 1.0
    message_count: 0.167
    token_estimate: 0.167
    code_blocks: 0.167
    tool_calls: 0.167
    math_expressions: 0.167
    reasoning_keywords: 0.167

  composite_weights:                 # Must sum to 1.0
    heuristic: 0.4
    ml: 0.6

  online_optimizer:
    enabled: false
    alpha: 0.01
    interval_secs: 3600
    state_path: null
    quality_threshold: 0.7

  quality_evaluator:
    enabled: false
    threshold: 0.3

  semantic_cache:
    enabled: false
    similarity_threshold: 0.95
    max_entries: 10000
    ttl_secs: 3600
    embedding_model: builtin-minilm
    min_quality_score: 0.7

  budget_limits: {}                  # Model group name → {hourly_limit_usd, daily_limit_usd, monthly_limit_usd}
  ab_test: null
  training:
    dataset_path: null
    learning_rate: 0.001
    batch_size: 32
    epochs: 10
    augmentation: false

  model_group_overrides: {}

Task Specializations

Models can declare task categories they excel at:

models:
  - provider: openai
    model: gpt-4.1
    tier: powerful
    specializations: [code_generation, factual_qa]

Available task types allow the router to prefer models with matching specializations when the request is classified into a particular task category.


Safe Routing Simulation

The Admin Panel includes a Safe Routing Simulation tool that lets you test routing decisions without sending a request to any provider:

  1. Enter a Model Group name (must match a configured group with tiered models)
  2. Enter a Prompt (the user message content to classify)
  3. Click "Simulate Without Generation"

The simulator runs the full classification pipeline (heuristic scoring, tier selection, context filtering, budget check) and returns:

  • Complexity score and task type
  • Selected tier and model
  • Candidates excluded for context capacity
  • Whether cascade would trigger

This is useful for:

  • Validating tier boundaries before enabling smart routing in production
  • Testing custom heuristic weights against real prompts
  • Verifying that expensive models are only selected for genuinely complex requests

ONNX Model Assets

The Smart Routing tab includes an ONNX Model Assets section for managing the ML classifier model:

Action Purpose
Check Assets Validates the artifact directory: checks for manifest.json, model.onnx, tokenizer.json, and .onnxruntime/
Install Approved ONNX Assets Downloads and installs the pinned, approved classifier model and runtime into the configured directory

The assets are installed server-side (not in the browser). Docker users should mount a writable persistent volume at the artifact directory path.


Next Steps

Clone this wiki locally