-
Notifications
You must be signed in to change notification settings - Fork 0
Reasoning Compatibility
Reasoning models (OpenAI o-series, Anthropic extended thinking, DeepSeek-R1, Qwen QwQ, and others) each carry reasoning state in their own shape — thinking blocks, redacted_thinking, reasoning_content, or reasoning fields. When a request fails over from one model family to another, replaying the previous model's reasoning state can produce 400 errors or corrupt the target model's context.
The reasoning-compatibility layer makes cross-model failover safe. It detects reasoning state in prior assistant turns, strips or preserves it based on the source and target model families, normalizes reasoning parameters to the target's accepted shape, and attributes reasoning-token spend in logs, metrics, and cost calculations.
The layer runs per failover attempt on a cloned outgoing request. It is enabled by default. Setting reasoning_compat.enabled: false reproduces exact passthrough behavior — no detection, no strip, no normalization, no cost changes.
The layer is a four-stage pipeline that runs on each failover attempt:
detect → policy → normalize → cost
-
Detect — Scans prior assistant turns for reasoning carriers: Anthropic
thinkingandredacted_thinkingblocks, OpenAI/DeepSeekreasoning_content, and genericreasoningfields. - Policy — Decides whether to strip or preserve the detected reasoning state, keyed off the source model family versus the target model family.
-
Normalize — Converts reasoning parameters to the shape the target model family accepts (for example, OpenAI-style
reasoning_effortto Anthropic manual-modethinking.budget_tokens). - Cost — Attributes reasoning-token usage (including invisible reasoning tokens) to the cost pipeline.
All actions are counts-only in logs: the layer never logs reasoning text, signatures, or redacted payloads.
The default behavior strips reasoning state whenever the failover target model differs from the model that produced it (strip_on_model_change: true). This is the safe default — a target model never receives another family's reasoning carriers.
When strip_on_model_change is false, reasoning state is preserved when the source and target share a reasoning family (for example, two Anthropic manual-mode models), and stripped only when the families differ.
conversation_model_affinity (default true) tracks which model produced each conversation prefix so the policy stage has accurate source-model attribution when making preserve decisions.
OpenAI-style reasoning_effort levels are mapped to Anthropic manual-mode thinking.budget_tokens using the effort-budget map. The defaults are:
| Effort | Budget (tokens) |
|---|---|
| minimal | 1024 |
| low | 2048 |
| medium | 8192 |
| high | 16384 |
| xhigh | 32768 |
Every budget must be at least 1024 (the Anthropic extended-thinking floor). A custom effort_budget_map must define all five efforts.
Reasoning Compatibility is configured under the Routing Settings tab, alongside Retry and Cache-Aware Routing.

The panel exposes:
-
Enable Reasoning Compatibility — master switch (
enabled). -
Strip Reasoning State on Model Change —
strip_on_model_change. -
Attribute Reasoning-Token Costs —
attribute_reasoning_cost. -
Track Conversation Model Affinity —
conversation_model_affinity. - Effort budgets — Minimal, Low, Medium, High, XHigh (all ≥ 1024).
- Per-Provider Overrides — override the effort map, strip behavior, or normalization per configured provider.
- Recent Reasoning Events — recent strip/normalize decisions by trace ID (counts and model IDs only).
reasoning_compat:
enabled: true # Master switch (default: true)
strip_on_model_change: true # Strip reasoning state on any model change
attribute_reasoning_cost: true # Attribute reasoning-token spend in cost/metrics
conversation_model_affinity: true # Track prefix → resolved model for attribution
effort_budget_map: # reasoning_effort → thinking.budget_tokens
minimal: 1024
low: 2048
medium: 8192
high: 16384
xhigh: 32768
per_provider: # Optional per-provider overrides
anthropic:
strip_on_model_change: false
effort_budget_map:
minimal: 1024
low: 4096
medium: 12288
high: 24576
xhigh: 32768Keys under per_provider must match configured provider names. Each override may set:
-
effort_budget_map— replace the global effort map for this provider (same ≥ 1024 rule, all five efforts required). -
strip_on_model_change— override the global strip behavior. -
normalize_reasoning_parameters— override reasoning-parameter normalization.
- Every budget in the effort map must be ≥ 1024.
- A custom effort map must define all five efforts (minimal, low, medium, high, xhigh).
-
per_providerkeys must reference known provider names. - Per-provider override effort maps follow the same budget rules.
Invalid configuration is rejected on load and on hot reload.
- Routing & Failover — cross-provider failover that this layer protects
- OAuth & Codex — reasoning-effort allowlists for Codex
- Configuration — full config reference