Skip to content

Cache Aware Routing

fdanobey edited this page Sep 2, 2026 · 1 revision

Cache-Aware Routing

Many providers offer prompt caching — reusing a previously processed prompt prefix at a fraction of the input-token price. Cache hits only happen when the same conversation prefix reaches the same provider, and (for explicit-cache providers like Anthropic Claude) when cache_control breakpoints mark the reusable prefix.

Cache-aware routing keeps conversations aligned with the provider that already holds their cached prefix, injects cache breakpoints for providers that need them, and factors cached-token pricing into cost-based provider selection.

The feature is disabled by default. Enable it with cache_aware_routing.enabled: true.

Cache-aware routing is distinct from the Token Compression cache-aware downgrade, which preserves cached prefixes byte-for-byte while compressing the suffix. Cache-aware routing operates on provider selection and cache-breakpoint injection.


What It Does

1. Sticky provider selection

The gateway computes a stable hash of the conversation prefix (system prompt, tools, and all but the trailing turn — the model name is part of the hash). On a successful response it records prefix_hash → (provider, model) in a short-lived sticky cache. Subsequent requests with the same prefix are routed back to the same provider so its cache stays warm.

Stickiness is a soft preference, not a pin:

  • If the sticky provider's circuit breaker is open, the entry is skipped (but left in place, so stickiness resumes once the breaker closes) and normal priority/cost/latency routing serves the request.
  • If the sticky provider is no longer a candidate for the model group, the entry is ignored.
  • Promotion moves the sticky provider to the front of the candidate list; it never removes other candidates.

2. Cache-breakpoint injection

For providers that advertise explicit prompt-cache support (a cache_support with a max_breakpoints limit), the gateway injects gateway-computed cache_control breakpoints into the outgoing request, honoring the per-model cache_min_tokens (or default_cache_min_tokens when the model has no explicit value) and the provider's max_breakpoints.

For OpenRouter, the gateway attaches a deterministic session_id derived from the prefix hash (obey-<hash>) so the intermediary's own stickiness aligns with the gateway's.

3. Cache-aware cost sorting

When cost_sort_hit_rate > 0.0, the routing cost comparison blends each candidate's cache-read input price with its uncached input price at the assumed hit rate. A candidate that is cheaper on cache reads sorts ahead of one that is only cheaper uncached. A hit rate of 0.0 keeps cost sorting identical to the uncached behavior.


Admin Panel

Cache-Aware Routing is configured under the Routing Settings tab.

Routing Settings — Cache-Aware Routing

The panel exposes:

  • Enable Cache-Aware Routing — master switch (enabled).
  • Stickiness TTL (seconds) — how long prefix→provider affinity entries live. 0 disables stickiness even when the feature is enabled.
  • Default Cache Min Tokens — fallback minimum uncached prefix size eligible for caching when a model has no explicit cache_min_tokens.
  • Cost Sort Hit Rate — assumed cache hit rate (0–1) used when weighting cache-read versus uncached input prices during cost sorting.

Configuration

cache_aware_routing:
  enabled: false                 # Master switch (default: false)
  stickiness_ttl_seconds: 300    # TTL for prefix → provider affinity (0 disables stickiness)
  default_cache_min_tokens: 1024 # Fallback min prefix tokens eligible for caching
  cost_sort_hit_rate: 0.0        # Assumed hit rate (0.0–1.0) for cache-aware cost sort

Per-model cache support

The breakpoint injector and cost sort read prompt-cache metadata from each model entry:

model_groups:
  - name: "claude-group"
    models:
      - provider: "anthropic"
        model: "claude-sonnet-4"
        cost_per_million_input_tokens: 3.0
        cost_per_million_output_tokens: 15.0
        # Prompt-cache metadata used by cache-aware routing:
        cache_min_tokens: 1024
        # cache_support / cache-read pricing are provider-model attributes

Interaction with Reasoning Compatibility

The prefix→provider sticky entries are shared with the Reasoning Compatibility conversation-model-affinity feature. Because both ride the same entries, stickiness_ttl_seconds must be greater than 0 for either to record affinity. When cache-aware routing is disabled but reasoning-compat affinity is enabled, the sticky cache still runs with the configured TTL to supply source-model attribution.


Notes

  • Stickiness never overrides an open circuit breaker or a non-candidate provider.
  • Breakpoint injection is a no-op unless cache-aware routing is enabled and the target model declares explicit cache support.
  • If breakpoint injection fails, the request is sent without cache markers rather than failing.

Related

Clone this wiki locally