You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Fix reasoning models silently failing on OpenAI-compatible servers
Real failure from testing: a local reasoning model (Qwen3-family, served
via llama.cpp based on the response's "timings" field) burned its entire
8192-token budget writing to a separate "reasoning_content" field, hit
finish_reason:"length", and left "content" empty - which just produced a
generic, unhelpful "OpenAI-compatible server returned no content".
- call_openai() now checks for this specific signature (empty content +
finish_reason "length" + non-empty reasoning_content) and reports
exactly what happened instead of a generic failure.
- Gives openai its own completion budget (openai_budget, default 16384,
separate from the shared $budget Anthropic/Gemini use) rather than the
8192 that was exhausted here - reasoning-capable models are common on
exactly the kind of local/self-hosted server this provider is also
meant to cover, and their reasoning traces alone can be substantial on
a prompt this size.
Verified both against the user's actual failure response (correctly
detected) and a normal successful response (unaffected, no false
positive).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>