Repository navigation
v2.57.1
Forward Qwen reasoning levels for chat-template models
Custom OpenAI-compatible Qwen models using thinkingFormat: "qwen-chat-template" no longer collapse every enabled dreb reasoning level to a single enable_thinking boolean.
What changed
- Canonical payload. The
qwen-chat-templateformat now sends nestedchat_template_kwargs.enable_thinkingplus the mapped effort as top-levelreasoning_effort(the canonical API used by the official Qwen3.8 example and llama.cpp); effort is omitted when thinking is off. - Qwen3.8+ vocabulary map. Qwen3.8-and-later models get a default mapping onto the three native tiers —
minimal/low→low,medium→medium,high/xhigh→xhigh— with any explicit per-modelreasoningEffortMapalways taking precedence. xhighselectable for Qwen3.8+.supportsXhigh()recognizes Qwen3.8+ ids so thexhighselection is preserved end-to-end instead of clamped.- Hardened version parsing.
qwenFamilyVersionrejects parameter-count digits (qwen-14b,Qwen-7B-Chat,deepseek-r1-distill-qwen-32b), embedded substrings, and two-digit majors, so only genuine Qwen family versions trigger the 3.8+ behavior; pre-3.8 Qwen models are unchanged. - Regression coverage. Wire-level payload tests through the real
streamSimplepath cover all five dreb labels, the explicit-map override, disabled reasoning, and negative cases proving olderqwen-chat-templatemodels do not inherit Qwen3.8 semantics. - Documentation updates.
packages/aiREADME, models guide, custom-provider guide, andOpenAICompletionsCompatJSDoc describe the canonical payload, three native tiers, and a concrete Qwen3.8 example config.
Implemented in PR 466.