-
-
Notifications
You must be signed in to change notification settings - Fork 7.5k
Thinking Budget
Dashboard: Settings → AI → Thinking Budget
API:GET/PUT/api/settings/thinking-budget
Source:open-sse/services/thinkingBudget.ts
Thinking Budget controls whether OmniRoute rewrites client thinking/reasoning parameters on the way to providers. It does not turn compression, routing, or prompt cache on or off.
| Mode | What OmniRoute does | When to use |
|---|---|---|
passthrough (default) |
Leaves client fields alone (reasoning, reasoning_effort, Claude thinking, Gemini thinking_config, etc.). |
Codex / Desktop / any client that should control effort + reasoning summaries. Required for visible thinking panels when the client requests reasoning.summary. |
auto |
Strips all thinking/reasoning fields from the request body before upstream. | Only when you deliberately want the provider to invent defaults and you do not need client-controlled thinking. Not “auto-show thinking”. |
custom |
Overwrites every request with a fixed thinking token budget. | Hard cap on thinking tokens for all traffic. |
adaptive |
Scales budget from a base effort using message count, tools, and prompt length. | Soft token control without fully stripping client intent. |
When mode is auto, stripThinkingConfig() deletes (among others):
- OpenAI / Responses:
reasoning,reasoning_effort - Claude:
thinking, andoutput_config.effortwhen present - Gemini:
generationConfig.thinking_config/thinkingConfig
If a client (e.g. Codex Desktop) sent reasoning: { effort: "ultra", summary: "detailed" }, auto drops that object. Upstream may still bill some reasoning tokens, but often returns empty or encrypted-only reasoning items — so the UI shows no useful thinking stream.
| Feature | Relationship |
|---|---|
| Compression (Caveman, RTK, stacked, …) | Separate pipeline. Works under every thinking-budget mode. |
| Prompt / semantic cache | Separate. Unaffected by thinking-budget mode. |
| Combo routing / fallbacks | Separate. Unaffected. |
| API-key token limits / cost budgets | Separate. Unaffected. |
| Reasoning replay cache | Multi-turn re-inject for strict providers (DeepSeek, Kimi, Qwen-thinking, …). Not the same as Desktop “show thinking”. |
Decrypting encrypted_content |
Impossible. OpenAI/Codex private reasoning blobs are opaque. OmniRoute never decrypts them (#7095 / #7176 / #7304). |
For a client to show thinking text you need all of:
- Thinking Budget mode =
passthrough(or custom/adaptive that still leaves summary requests intact enough for the path you use). - Client asks for a summary, e.g. Codex
model_reasoning_summary = "detailed"/auto(notnone). - Upstream actually streams
response.reasoning_summary_text.*(or a non-emptyreasoning.summaryon the item).
If you only get “encrypted private reasoning”, either:
- mode was
auto(client request was stripped), or - upstream returned
encrypted_contentwithout summary text (provider limitation; OmniRoute can only surface a placeholder, not plaintext).
# Read
curl -sS https://localhost:20128/api/settings/thinking-budget \
-H "Authorization: Bearer $OMNIROUTE_TOKEN"
# Recommended for Codex / Desktop thinking visibility
curl -sS -X PUT https://localhost:20128/api/settings/thinking-budget \
-H "Authorization: Bearer $OMNIROUTE_TOKEN" \
-H "Content-Type: application/json" \
-d '{"mode":"passthrough","customBudget":10240,"effortLevel":"medium"}'Schema (updateThinkingBudgetSchema): mode ∈ passthrough|auto|custom|adaptive; optional customBudget, effortLevel, baseBudget, complexityMultiplier.
Value is stored under settings key thinkingBudget and hydrated at process start (hydrateThinkingBudgetConfig). After changing via DB or some non-API paths, restart the OmniRoute process so the in-memory singleton matches disk.
- Codex / Desktop users: mode = passthrough
- Compression still enabled if you want token savings on messages, not by stripping thinking
- Do not expect
autoto “show more thinking” - Encrypted-only summaries are a provider behavior; passthrough cannot decrypt them
-
REASONING_REPLAY.md — multi-turn
reasoning_contentcache - USER_GUIDE.md — Settings dashboard tabs
- API_REFERENCE.md — settings endpoints
OmniRoute · Website · npm · Docker Hub
- Setup Guide
- User Guide
- Features
- Quick Start (Docker)
- Electron Desktop App
- Termux (Android)
- PWA Guide
- MCP Server
- A2A Server
- Agent Protocols
- OpenCode Plugin
- Webhooks
- Cloud Agents
- Skills
- Memory
- Evals
- Gamification
- Guardrails
- Compliance
- Error Sanitization
- Public Credentials
- Route Guard Tiers
- Stealth Guide
- CLI Token Auth