-
Notifications
You must be signed in to change notification settings - Fork 0
Routing and Failover
OBEY API Gateway's core strength is intelligent request routing with automatic failover. When a provider fails, the gateway transparently tries the next one — your application never notices.
-
Client sends a request with a model name (e.g.,
gpt-4-group) - Gateway resolves the model group matching that name
- Models within the group are sorted by priority (lower = higher priority)
-
Router checks each candidate in order:
- Is the circuit breaker closed (healthy)?
- Is the provider within its rate limit?
- What's the current latency?
- Request is forwarded to the first eligible provider
- On failure, the next provider in the fallback chain is tried
Client Request → Model Group Resolution → Priority Sorting
│
┌─────────────────────────┼────────────────────────┐
▼ ▼ ▼
Provider A Provider B Provider C
(priority: 1) (priority: 2) (priority: 3)
┌─────┐ ┌─────┐ ┌─────┐
│ CB │ │ CB │ │ CB │
│Check│ │Check│ │Check│
└──┬──┘ └──┬──┘ └──┬──┘
│ │ │
[Success] ──────────────────────────────────────▶ Response
│
[Failure] ──▶ Try Provider B ──▶ Try Provider C ──▶ Error
Model groups define equivalent models that can serve the same requests:
model_groups:
- name: "gpt-4-group"
version_fallback_enabled: false
models:
- provider: "openai"
model: "gpt-4"
cost_per_million_input_tokens: 10.0
cost_per_million_output_tokens: 30.0
priority: 1 # Highest priority (tried first)
- provider: "ollama-local"
model: "llama3"
priority: 2 # Fallback
- provider: "together"
model: "meta-llama/Meta-Llama-3-70B"
priority: 3 # Last resortClients reference the group name as the model:
client.chat.completions.create(model="gpt-4-group", ...)
When version_fallback_enabled: true, the gateway sorts the whole group by version date (newest first) — the newest dated model is preferred even over lower-priority or cheaper models, undated models keep their normal priority/cost/latency order after all dated ones, and failover walks older versions before undated fallbacks. Useful for models that frequently update.
Each provider has an independent circuit breaker that prevents repeated calls to failing providers.
circuit_breaker:
failure_threshold: 3 # Consecutive failures to trip
backoff_sequence_seconds: [5, 10, 20, 40, 300]
success_threshold: 1 # Successes to recover
| State | Behavior |
|---|---|
| Closed | Normal operation — requests flow through |
| Open | Provider is marked unhealthy — requests skip immediately |
| Half-Open | After backoff period — one test request allowed |
After tripping, the circuit breaker waits progressively longer between recovery attempts:
Trip → 5s wait → test → fail → 10s wait → test → fail → 20s wait → ...
The sequence resets on successful recovery.
All circuit breakers are cleared when configuration is reloaded via the admin API. This lets you manually recover a provider after fixing the underlying issue.
The gateway detects rate limiting across multiple signal types and instantly fails over:
| Signal | Detection |
|---|---|
| HTTP 429 | Standard rate limit response |
| Rate-limit-shaped 200 | Response body contains rate limit error despite 200 status |
Retry-After header |
Honored for cooldown duration |
X-RateLimit-Reset header |
Used to calculate cooldown |
| Anthropic ISO reset headers | Parsed for exact reset time |
| Weekly-quota providers | Per-provider cooldown overrides for providers like Nano-GPT |
When rate-limited, the provider is temporarily skipped (not circuit-broken — rate-limit failures do not count toward the circuit-breaker threshold) and requests route to the next provider immediately. The cooldown also applies to streaming pass-through attempts, which consume the provider's rate_limit_per_minute tokens just like non-streaming requests.
Failed requests are retried within a provider before failing over to the next:
retry:
max_retries_per_provider: 1 # Retries per provider
backoff_sequence_seconds: [1, 2, 4]Retries use exponential backoff and only trigger for transient errors (5xx, timeouts, network errors). Permanent errors (4xx) fail immediately.
Models within a group are sorted by priority (lower = higher priority). For equal-priority models, the gateway considers:
-
Cost —
cost_per_million_input_tokensandcost_per_million_output_tokensfor budget-aware selection - Latency — tracked per-provider for optimal response times
- Provider health — circuit breaker and rate limit status
models:
- provider: "openai"
model: "gpt-4"
cost_per_million_input_tokens: 10.0
cost_per_million_output_tokens: 30.0
priority: 1 # Primary
- provider: "together"
model: "meta-llama/Meta-Llama-3-70B"
cost_per_million_input_tokens: 0.9
cost_per_million_output_tokens: 0.9
priority: 2 # Cheaper fallbackWhen providers support prompt caching, the gateway can keep a conversation aligned with the provider that already holds its cached prefix, inject cache_control breakpoints for explicit-cache providers, and factor cached-token pricing into the equal-priority cost sort. This is configured under the Routing Settings tab alongside Retry and Reasoning Compatibility.

See Cache-Aware Routing for full details.
Failing a reasoning model over to a different model family can corrupt context or trigger 400 errors if the previous model's reasoning state (thinking blocks, reasoning_content, and similar) is replayed as-is. The gateway's Reasoning Compatibility layer detects that state, strips or preserves it based on the source and target families, and normalizes reasoning parameters for the target — enabled by default so cross-model failover stays safe.
When a request exceeds a model's context window limit, the gateway automatically truncates messages to fit rather than failing with a context-length error. This ensures requests always have a chance of succeeding even with large conversation histories.
The gateway maintains rolling latency statistics per provider. This data is:
- Used for routing decisions between equal-priority providers
- Displayed in the dashboard metrics
- Available via the Prometheus endpoint

When all providers in a model group fail, the gateway returns an aggregated error showing each attempt:
{
"error": {
"message": "All providers failed",
"type": "gateway_error",
"attempts": [
{"provider": "openai", "error": "timeout after 30s"},
{"provider": "ollama", "error": "connection refused"},
{"provider": "together", "error": "rate limited (retry after 60s)"}
]
}
}- Streaming — how streaming works with failover
- Providers — configure your providers
- Configuration — full config reference