Which version of LM Studio?
LM Studio 0.3.22-1 (Linux)
Which operating system?
Ubuntu 24.04.2 LTS
What is the bug?
When calling the HTTP API (/v1/chat/completions) with a "Thinking" variant of Qwen3 (e.g., qwen3-4b-thinking-2507), the <think>...</think> block is missing from the response.
This used to work correctly in previous versions of LM Studio — the block was included in the API response and could be parsed by my local application. After updating to 0.3.22, the content is no longer included in the API output, even though:
- The LM Studio GUI still shows the internal Thinking blocks, indicating the model is generating them.
- I did not change the Jinja template or any model settings.
- Resetting all LM Studio settings had no effect.
- The issue occurs with multiple Qwen3-thinking models, not just one.
To Reproduce
curl http://localhost:1234/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-4b-thinking-2507",
"messages": [
{ "role": "system", "content": "Always answer in rhymes. Today is Thursday" },
{ "role": "user", "content": "What day is it today?" }
],
"temperature": 0.7,
"max_tokens": -1,
"stream": false
}'
Example response:
[qwen3-4b-thinking-2507] Generated prediction:
{
"id": "chatcmpl-evf5zdsvuaonsed2l41gw",
"object": "chat.completion",
"created": 1754495472,
"model": "qwen3-4b-thinking-2507",
"choices": [
{
"index": 0,
"logprobs": null,
"finish_reason": "stop",
"message": {
"role": "assistant",
"content": "\nOkay, the user asked \"What day is it today?\" but also specified that I should always respond in rhymes and that today is Thursday. \n\nHmm, this is interesting—they're testing if I follow instructions creatively while staying accurate. They probably want something playful but precise since they mentioned Thursdays specifically. \n\nI recall they emphasized \"always answer in rhymes,\" so no dry facts allowed. Gotta craft a short, cheerful couplet that confirms Thursday without sounding robotic. \n\n*Brainstorming lines...* \nFirst line should state the day clearly: \"Today's the fourth day of the week\" → too wordy? Wait, \"Thursday\" is simpler! \nSecond line needs rhyme and lightness... maybe add why it matters (\"best for plans\" feels universally relatable). \n\nUser didn't give context about their mood, but since they're asking for rhymes, they likely want fun. No need to overcomplicate—just a tiny cheerful nudge. \n\n*Double-checking*: \n- Rhyme? \"Day/plan\" works (slant rhyme, but friendly enough). \n- Thursday confirmed? Yep. \n- Short? Under 10 words per line → good. \n\n...And no markdown! Pure poetry in their request. *Nods*. Done.\n\n\nToday's the day that starts with \"Th\" and ends with \"day\"— \n**Thursday**, bright and full of plans to make it play! 🌟"
}
}
],
"usage": {
"prompt_tokens": 28,
"completion_tokens": 295,
"total_tokens": 323
},
"stats": {},
"system_fingerprint": "qwen3-4b-thinking-2507"
}
Expected behavior:
The response should include a think-block with internal reasoning (as it did in earlier versions).
Which version of LM Studio?
LM Studio 0.3.22-1 (Linux)
Which operating system?
Ubuntu 24.04.2 LTS
What is the bug?
When calling the HTTP API (/v1/chat/completions) with a "Thinking" variant of Qwen3 (e.g., qwen3-4b-thinking-2507), the <think>...</think> block is missing from the response.
This used to work correctly in previous versions of LM Studio — the block was included in the API response and could be parsed by my local application. After updating to 0.3.22, the content is no longer included in the API output, even though:
To Reproduce
Example response:
[qwen3-4b-thinking-2507] Generated prediction:
Expected behavior:
The response should include a think-block with internal reasoning (as it did in earlier versions).