Skip to content

Smart Routing

fdanobey edited this page Aug 7, 2026 · 4 revisions

Smart Model Routing

Smart Model Routing automatically selects the best model within a model group based on request complexity, cost constraints, and quality feedback. Instead of always routing to the highest-priority provider, the system classifies each request's complexity and dispatches it to the most cost-effective tier that can handle it well.

Smart routing is opt-in (enabled: false by default) and operates alongside the existing priority-based failover. When disabled, routing falls back to standard priority ordering.


How It Works

Request ──► Complexity Classifier ──► Tier Selection ──► Model Dispatch
                    │                       │
            ┌───────┼───────┐       ┌───────┼───────┐
            │       │       │       │       │       │
         Heuristic  ML    LLM    Fast   Balanced  Powerful
            │       │       │       │       │       │
            └───────┼───────┘       └───────┼───────┘
                    │                       │
                    ▼                       ▼
            Complexity Score         Best-fit Model
              (0.0–1.0)           (cost + quality)
  1. Classify — The request is scored for complexity (0.0 = trivial, 1.0 = highly complex)
  2. Map to tier — The score is mapped to a capability tier using configurable boundaries
  3. Select model — Within the tier, the most cost-effective model with available capacity is chosen
  4. Cascade (optional) — If the response quality is insufficient, escalate to a higher tier

Model Tiers

Each model in a model group can be assigned a capability tier:

Tier Use Case Example Models
Fast Simple queries, lookups, classification GPT-3.5, Llama 3 8B, Gemini Flash
Balanced Multi-step reasoning, code generation GPT-4o, Claude Sonnet, Llama 3 70B
Powerful Complex analysis, research, long-context GPT-4.1, Claude Opus, o3
model_groups:
  - name: "smart-group"
    models:
      - provider: "openai"
        model: "gpt-4.1"
        tier: powerful
        context_window: 128000
        specializations: [code_generation, factual_qa]
        priority: 1

      - provider: "openai"
        model: "gpt-4o"
        tier: balanced
        context_window: 128000
        priority: 2

      - provider: "groq"
        model: "llama3-8b-8192"
        tier: fast
        context_window: 8192
        priority: 3

Tier Boundaries

The complexity score (0.0–1.0) maps to tiers via configurable boundaries:

smart_routing:
  tier_boundaries:
    fast_max: 0.3        # Scores 0.0–0.3 → Fast tier
    balanced_max: 0.7    # Scores 0.3–0.7 → Balanced tier
                         # Scores 0.7–1.0 → Powerful tier

Classifier Modes

The gateway offers four complexity classification strategies:

Mode Description Requirements
heuristic (default) Weighted signal analysis of request structure None
ml ONNX machine-learning model inference Trained model file
llm Uses a configured LLM to assess complexity Provider + model
composite Weighted blend of heuristic + ML scores Heuristic + ML model

Heuristic Classifier

Analyzes request signals with configurable weights (must sum to 1.0):

smart_routing:
  classifier: heuristic
  heuristic_weights:
    message_count: 0.167       # Number of conversation messages
    token_estimate: 0.167      # Estimated input token count
    code_blocks: 0.167         # Presence of fenced code blocks
    tool_calls: 0.167          # Number of tool definitions
    math_expressions: 0.167    # Mathematical notation detected
    reasoning_keywords: 0.167  # Keywords indicating complex reasoning

ML Classifier

Uses a pre-trained ONNX model for complexity scoring:

smart_routing:
  classifier: ml
  ml_model_path: ./models/complexity_classifier.onnx

LLM Classifier

Delegates complexity assessment to a configured model:

smart_routing:
  classifier: llm
  classifier_model: gpt-4o-mini

Composite Classifier

Blends heuristic and ML scores:

smart_routing:
  classifier: composite
  ml_model_path: ./models/complexity_classifier.onnx
  composite_weights:
    heuristic: 0.4
    ml: 0.6

Cascade Escalation

When enabled, cascade monitors the response quality from a lower tier and escalates to a higher tier if the response is insufficient:

smart_routing:
  cascade:
    enabled: true
    max_escalations: 2          # Maximum tier jumps (1–2)
    min_response_tokens: 50     # Minimum tokens before evaluating quality
    early_signal_tokens: 100    # Tokens to inspect for early escalation signals

Streaming Cascade Mode

Mode Behavior
buffer (default) Buffer the response, evaluate quality, re-dispatch if needed
early_signal Inspect the first N tokens during streaming; escalate before completion
smart_routing:
  streaming_cascade_mode: buffer   # buffer | early_signal

Online Optimizer

The online optimizer continuously adjusts tier boundaries based on observed quality scores:

smart_routing:
  online_optimizer:
    enabled: true
    alpha: 0.1                  # EMA learning rate (0.0–1.0)
    interval_secs: 3600         # Optimization interval (1–604800)
    state_path: ./smart_routing_state.json
    quality_threshold: 0.7      # Minimum acceptable quality (0.0–1.0)

When quality drops below the threshold for a tier, the optimizer shifts boundaries to route more requests to higher tiers. When quality exceeds the threshold consistently, boundaries relax to save cost.


Quality Evaluator

Evaluates response quality to feed the cascade and optimizer:

smart_routing:
  quality_evaluator:
    enabled: true
    threshold: 0.7              # Quality score below this triggers escalation

Semantic Cache

Smart routing includes its own semantic cache (separate from the gateway's response cache) that caches complexity classification decisions:

smart_routing:
  semantic_cache:
    enabled: true
    similarity_threshold: 0.9
    max_entries: 10000
    ttl_secs: 3600
    embedding_model: builtin-minilm
    min_quality_score: 0.7

When a new request is semantically similar to a previously classified request, the cached tier decision is reused — skipping classification entirely.


Budget Limits

Per-model-group spend limits for smart routing decisions:

smart_routing:
  budget_limits:
    smart-group:
      hourly_limit_usd: 5.0
      daily_limit_usd: 50.0
      monthly_limit_usd: 500.0

When a budget limit is reached, the router favors cheaper tiers or falls back to standard priority routing.


A/B Testing

Run controlled experiments comparing two routing policies:

smart_routing:
  ab_test:
    variant_percentage: 0.2     # 20% of traffic uses the variant policy
    control:
      classifier: heuristic
      cost_quality_threshold: 0.5
    variant:
      classifier: composite
      ml_model_path: ./models/new_classifier.onnx
      cost_quality_threshold: 0.7

The variant_percentage (0.0–1.0) controls what fraction of requests use the variant policy. Quality metrics are tracked separately for each arm.


Training

Configure offline training jobs for the ML classifier:

smart_routing:
  training:
    dataset_path: ./training_data.jsonl
    learning_rate: 0.001
    batch_size: 32
    epochs: 10
    augmentation: true

Training data is collected from the online optimizer's quality observations. The trained model replaces the existing ml_model_path after validation.


Per-Model-Group Overrides

Each model group can override any global smart routing setting:

smart_routing:
  enabled: true
  classifier: heuristic

  model_group_overrides:
    coding-group:
      classifier: composite
      ml_model_path: ./models/code_classifier.onnx
      cost_quality_threshold: 0.8
      cascade:
        enabled: true
        max_escalations: 1
    simple-group:
      enabled: false    # Disable smart routing for this group

Resolution: per-group override fields replace the corresponding global fields; omitted fields inherit from the global config.


Context Window Awareness

Smart routing considers each model's declared context_window when selecting targets:

  • Requests exceeding a model's context window skip that model
  • reserved_output_tokens (default: 4096) is subtracted from the available window
  • provider_overhead_tokens (default: 256) accounts for provider-injected content
  • context_safety_margin_tokens (default: 512) provides additional headroom
smart_routing:
  reserved_output_tokens: 4096
  provider_overhead_tokens: 256
  context_safety_margin_tokens: 512
  allow_unknown_context_window: false   # Reject models with context_window: 0

Configuration Reference

smart_routing:
  enabled: false
  classifier: heuristic              # heuristic | ml | llm | composite
  ml_model_path: null                # Path to ONNX classifier model
  classifier_model: null             # LLM model name (for llm classifier)
  cost_quality_threshold: 0.5        # Cost vs quality tradeoff (0.0–1.0)

  tier_boundaries:
    fast_max: 0.3
    balanced_max: 0.7

  cascade:
    enabled: false
    max_escalations: 2
    min_response_tokens: 50
    early_signal_tokens: 100

  streaming_cascade_mode: buffer     # buffer | early_signal

  heuristic_weights:                 # Must sum to 1.0
    message_count: 0.167
    token_estimate: 0.167
    code_blocks: 0.167
    tool_calls: 0.167
    math_expressions: 0.167
    reasoning_keywords: 0.167

  composite_weights:                 # Must sum to 1.0
    heuristic: 0.4
    ml: 0.6

  online_optimizer:
    enabled: false
    alpha: 0.1
    interval_secs: 3600
    state_path: null
    quality_threshold: 0.7

  quality_evaluator:
    enabled: false
    threshold: 0.7

  semantic_cache:
    enabled: false
    similarity_threshold: 0.9
    max_entries: 10000
    ttl_secs: 3600
    embedding_model: builtin-minilm
    min_quality_score: 0.7

  budget_limits: {}
  ab_test: null
  training:
    dataset_path: null
    learning_rate: 0.001
    batch_size: 32
    epochs: 10
    augmentation: false

  reserved_output_tokens: 4096
  provider_overhead_tokens: 256
  context_safety_margin_tokens: 512
  allow_unknown_context_window: false

  model_group_overrides: {}

Task Specializations

Models can declare task categories they excel at:

models:
  - provider: openai
    model: gpt-4.1
    tier: powerful
    specializations: [code_generation, factual_qa]

Available task types allow the router to prefer models with matching specializations when the request is classified into a particular task category.


Next Steps

Clone this wiki locally