Skip to content

v0.1.20

Choose a tag to compare

@github-actions github-actions released this 17 Sep 08:17
· 10 commits to master since this release
Fix reasoning models silently failing on OpenAI-compatible servers

Real failure from testing: a local reasoning model (Qwen3-family, served
via llama.cpp based on the response's "timings" field) burned its entire
8192-token budget writing to a separate "reasoning_content" field, hit
finish_reason:"length", and left "content" empty - which just produced a
generic, unhelpful "OpenAI-compatible server returned no content".

- call_openai() now checks for this specific signature (empty content +
  finish_reason "length" + non-empty reasoning_content) and reports
  exactly what happened instead of a generic failure.
- Gives openai its own completion budget (openai_budget, default 16384,
  separate from the shared $budget Anthropic/Gemini use) rather than the
  8192 that was exhausted here - reasoning-capable models are common on
  exactly the kind of local/self-hosted server this provider is also
  meant to cover, and their reasoning traces alone can be substantial on
  a prompt this size.

Verified both against the user's actual failure response (correctly
detected) and a normal successful response (unaffected, no false
positive).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>