Skip to content

v2.57.1

Choose a tag to compare

@aebrer aebrer released this 15 Aug 21:50
· 28 commits to master since this release
b83112b

Forward Qwen reasoning levels for chat-template models

Custom OpenAI-compatible Qwen models using thinkingFormat: "qwen-chat-template" no longer collapse every enabled dreb reasoning level to a single enable_thinking boolean.

What changed

  • Canonical payload. The qwen-chat-template format now sends nested chat_template_kwargs.enable_thinking plus the mapped effort as top-level reasoning_effort (the canonical API used by the official Qwen3.8 example and llama.cpp); effort is omitted when thinking is off.
  • Qwen3.8+ vocabulary map. Qwen3.8-and-later models get a default mapping onto the three native tiers — minimal/low → low, medium → medium, high/xhigh → xhigh — with any explicit per-model reasoningEffortMap always taking precedence.
  • xhigh selectable for Qwen3.8+. supportsXhigh() recognizes Qwen3.8+ ids so the xhigh selection is preserved end-to-end instead of clamped.
  • Hardened version parsing. qwenFamilyVersion rejects parameter-count digits (qwen-14b, Qwen-7B-Chat, deepseek-r1-distill-qwen-32b), embedded substrings, and two-digit majors, so only genuine Qwen family versions trigger the 3.8+ behavior; pre-3.8 Qwen models are unchanged.
  • Regression coverage. Wire-level payload tests through the real streamSimple path cover all five dreb labels, the explicit-map override, disabled reasoning, and negative cases proving older qwen-chat-template models do not inherit Qwen3.8 semantics.
  • Documentation updates. packages/ai README, models guide, custom-provider guide, and OpenAICompletionsCompat JSDoc describe the canonical payload, three native tiers, and a concrete Qwen3.8 example config.

Implemented in PR 466.