-
Notifications
You must be signed in to change notification settings - Fork 0
Cache Aware Routing
Many providers offer prompt caching — reusing a previously processed prompt prefix at a fraction of the input-token price. Cache hits only happen when the same conversation prefix reaches the same provider, and (for explicit-cache providers like Anthropic Claude) when cache_control breakpoints mark the reusable prefix.
Cache-aware routing keeps conversations aligned with the provider that already holds their cached prefix, injects cache breakpoints for providers that need them, and factors cached-token pricing into cost-based provider selection.
The feature is disabled by default. Enable it with cache_aware_routing.enabled: true.
Cache-aware routing is distinct from the Token Compression cache-aware downgrade, which preserves cached prefixes byte-for-byte while compressing the suffix. Cache-aware routing operates on provider selection and cache-breakpoint injection.
The gateway computes a stable hash of the conversation prefix (system prompt, tools, and all but the trailing turn — the model name is part of the hash). On a successful response it records prefix_hash → (provider, model) in a short-lived sticky cache. Subsequent requests with the same prefix are routed back to the same provider so its cache stays warm.
Stickiness is a soft preference, not a pin:
- If the sticky provider's circuit breaker is open, the entry is skipped (but left in place, so stickiness resumes once the breaker closes) and normal priority/cost/latency routing serves the request.
- If the sticky provider is no longer a candidate for the model group, the entry is ignored.
- Promotion moves the sticky provider to the front of the candidate list; it never removes other candidates.
For providers that advertise explicit prompt-cache support (a cache_support with a max_breakpoints limit), the gateway injects gateway-computed cache_control breakpoints into the outgoing request, honoring the per-model cache_min_tokens (or default_cache_min_tokens when the model has no explicit value) and the provider's max_breakpoints.
For OpenRouter, the gateway attaches a deterministic session_id derived from the prefix hash (obey-<hash>) so the intermediary's own stickiness aligns with the gateway's.
When cost_sort_hit_rate > 0.0, the routing cost comparison blends each candidate's cache-read input price with its uncached input price at the assumed hit rate. A candidate that is cheaper on cache reads sorts ahead of one that is only cheaper uncached. A hit rate of 0.0 keeps cost sorting identical to the uncached behavior.
Cache-Aware Routing is configured under the Routing Settings tab.

The panel exposes:
-
Enable Cache-Aware Routing — master switch (
enabled). -
Stickiness TTL (seconds) — how long prefix→provider affinity entries live.
0disables stickiness even when the feature is enabled. -
Default Cache Min Tokens — fallback minimum uncached prefix size eligible for caching when a model has no explicit
cache_min_tokens. - Cost Sort Hit Rate — assumed cache hit rate (0–1) used when weighting cache-read versus uncached input prices during cost sorting.
cache_aware_routing:
enabled: false # Master switch (default: false)
stickiness_ttl_seconds: 300 # TTL for prefix → provider affinity (0 disables stickiness)
default_cache_min_tokens: 1024 # Fallback min prefix tokens eligible for caching
cost_sort_hit_rate: 0.0 # Assumed hit rate (0.0–1.0) for cache-aware cost sortThe breakpoint injector and cost sort read prompt-cache metadata from each model entry:
model_groups:
- name: "claude-group"
models:
- provider: "anthropic"
model: "claude-sonnet-4"
cost_per_million_input_tokens: 3.0
cost_per_million_output_tokens: 15.0
# Prompt-cache metadata used by cache-aware routing:
cache_min_tokens: 1024
# cache_support / cache-read pricing are provider-model attributesThe prefix→provider sticky entries are shared with the Reasoning Compatibility conversation-model-affinity feature. Because both ride the same entries, stickiness_ttl_seconds must be greater than 0 for either to record affinity. When cache-aware routing is disabled but reasoning-compat affinity is enabled, the sticky cache still runs with the configured TTL to supply source-model attribution.
- Stickiness never overrides an open circuit breaker or a non-candidate provider.
- Breakpoint injection is a no-op unless cache-aware routing is enabled and the target model declares explicit cache support.
- If breakpoint injection fails, the request is sent without cache markers rather than failing.
- Routing & Failover — priority, cost, and latency routing this layer influences
- Token Compression — the separate cache-aware downgrade for compression
- Configuration — full config reference