Skip to content

Reasoning Compatibility

fdanobey edited this page Sep 2, 2026 · 1 revision

Reasoning Compatibility

Reasoning models (OpenAI o-series, Anthropic extended thinking, DeepSeek-R1, Qwen QwQ, and others) each carry reasoning state in their own shape — thinking blocks, redacted_thinking, reasoning_content, or reasoning fields. When a request fails over from one model family to another, replaying the previous model's reasoning state can produce 400 errors or corrupt the target model's context.

The reasoning-compatibility layer makes cross-model failover safe. It detects reasoning state in prior assistant turns, strips or preserves it based on the source and target model families, normalizes reasoning parameters to the target's accepted shape, and attributes reasoning-token spend in logs, metrics, and cost calculations.

The layer runs per failover attempt on a cloned outgoing request. It is enabled by default. Setting reasoning_compat.enabled: false reproduces exact passthrough behavior — no detection, no strip, no normalization, no cost changes.


How It Works

The layer is a four-stage pipeline that runs on each failover attempt:

detect → policy → normalize → cost
  1. Detect — Scans prior assistant turns for reasoning carriers: Anthropic thinking and redacted_thinking blocks, OpenAI/DeepSeek reasoning_content, and generic reasoning fields.
  2. Policy — Decides whether to strip or preserve the detected reasoning state, keyed off the source model family versus the target model family.
  3. Normalize — Converts reasoning parameters to the shape the target model family accepts (for example, OpenAI-style reasoning_effort to Anthropic manual-mode thinking.budget_tokens).
  4. Cost — Attributes reasoning-token usage (including invisible reasoning tokens) to the cost pipeline.

All actions are counts-only in logs: the layer never logs reasoning text, signatures, or redacted payloads.


Strip vs Preserve

The default behavior strips reasoning state whenever the failover target model differs from the model that produced it (strip_on_model_change: true). This is the safe default — a target model never receives another family's reasoning carriers.

When strip_on_model_change is false, reasoning state is preserved when the source and target share a reasoning family (for example, two Anthropic manual-mode models), and stripped only when the families differ.

conversation_model_affinity (default true) tracks which model produced each conversation prefix so the policy stage has accurate source-model attribution when making preserve decisions.


Effort-to-Budget Mapping

OpenAI-style reasoning_effort levels are mapped to Anthropic manual-mode thinking.budget_tokens using the effort-budget map. The defaults are:

Effort Budget (tokens)
minimal 1024
low 2048
medium 8192
high 16384
xhigh 32768

Every budget must be at least 1024 (the Anthropic extended-thinking floor). A custom effort_budget_map must define all five efforts.


Admin Panel

Reasoning Compatibility is configured under the Routing Settings tab, alongside Retry and Cache-Aware Routing.

Routing Settings — Reasoning Compatibility

The panel exposes:

  • Enable Reasoning Compatibility — master switch (enabled).
  • Strip Reasoning State on Model Change — strip_on_model_change.
  • Attribute Reasoning-Token Costs — attribute_reasoning_cost.
  • Track Conversation Model Affinity — conversation_model_affinity.
  • Effort budgets — Minimal, Low, Medium, High, XHigh (all ≥ 1024).
  • Per-Provider Overrides — override the effort map, strip behavior, or normalization per configured provider.
  • Recent Reasoning Events — recent strip/normalize decisions by trace ID (counts and model IDs only).

Configuration

reasoning_compat:
  enabled: true                       # Master switch (default: true)
  strip_on_model_change: true         # Strip reasoning state on any model change
  attribute_reasoning_cost: true      # Attribute reasoning-token spend in cost/metrics
  conversation_model_affinity: true   # Track prefix → resolved model for attribution
  effort_budget_map:                  # reasoning_effort → thinking.budget_tokens
    minimal: 1024
    low: 2048
    medium: 8192
    high: 16384
    xhigh: 32768
  per_provider:                       # Optional per-provider overrides
    anthropic:
      strip_on_model_change: false
      effort_budget_map:
        minimal: 1024
        low: 4096
        medium: 12288
        high: 24576
        xhigh: 32768

Per-Provider Overrides

Keys under per_provider must match configured provider names. Each override may set:

  • effort_budget_map — replace the global effort map for this provider (same ≥ 1024 rule, all five efforts required).
  • strip_on_model_change — override the global strip behavior.
  • normalize_reasoning_parameters — override reasoning-parameter normalization.

Validation

  • Every budget in the effort map must be ≥ 1024.
  • A custom effort map must define all five efforts (minimal, low, medium, high, xhigh).
  • per_provider keys must reference known provider names.
  • Per-provider override effort maps follow the same budget rules.

Invalid configuration is rejected on load and on hot reload.


Related

Clone this wiki locally