Replies: 3 comments
rc.1.2 复发数据点(2026-08-22,同签名,跨版本持续)环境:Windows 11 25H2 / Node.js 24.14.0 /
泄漏文本样例(存为标准 text 块,实为思考草稿;全英文流畅推理,与最终中文正式回复形成对照): 触发特征复现:多工具调用轮次 + 超长思考(本次为插件兼容调查轮,十余次工具调用、思考数千 token)——与首报及 §4.4 结论一致。 版本沿革:rc.7(首报)→ rc.1.1(重启后首轮未复现一次,样本 1)→ rc.1.2(复发,本数据点)。流式(reasoning-chunks=0)、存储(无 reasoning 块)、重放三层一致丢失,与首报机制相同。修复仍未进入任何发布版本。 |
|
+1 — reproduced on dsh 0.1.1-rc.2, Linux , Node 24 Same signature across all three layers (streaming / storage / replay): Route: commandcode / deepseek/deepseek-v4-pro / reasoningEffort: high One additional observation: in my case the leaked thinking text is entirely in a different language (Chinese thinking, English final reply), which makes the boundary very obvious — the text block starts with thousands of characters of Chinese reasoning followed by a short English answer at the tail. This is a blocking issue for any reasoning model used through tool-heavy workflows. Looking forward to a fix in the BlockAssembler / ReplayEnvelope seam. |
|
I expanded an independent troubleshooting guide with this distinction: https://github.com/sandbaseai/deepseek-harness-handbook/blob/main/docs/en/troubleshooting/response-language-and-reasoning.md It separates language preference from an assembler/channel-classification defect, and recommends comparing durable block shape plus chunk counts before changing prompts. The guide treats leaked reasoning as a privacy and output-budget issue until a release fixes the seam. |
Uh oh!
There was an error while loading. Please reload this page.
[Bug] rc.7: reasoning (thinking) intermittently stored and rendered as text blocks — thinking leaks into the visible transcript
Category: General (the repo has no Bug Report category; following the
[Bug]title convention of #274)Reporter environment: dsh
0.1.0-rc.7(npm, commit99f6f02fec), Web GUI athttp://127.0.0.1:3080, Windows 11 x64, Node 241. Summary
On dsh 0.1.0-rc.7, in the Web GUI, some assistant turns have their reasoning content stored and rendered as a regular
textblock instead of areasoningblock. The thinking that should be folded under the "Think" disclosure appears inline in the transcript as normal text. It is intermittent: healthy turns dominate, but affected turns leak — sometimes entirely (the turn'sreasoning-chunksevent count drops to 0 and the thinking is projected into thetextchannel).The misclassification is consistent across all three layers we inspected: the streaming chunk projection, the final stored
assistant/message, and the replay metadata. The client rendering layer is ruled out (byte-identical reasoning dispatch in rc.6 vs rc.7 bundles). The assembly layer (BlockAssembler / the pi-ai adapter path) is the suspected seam.2. Environment
0.1.0-rc.7(npmlatest, commit99f6f02fec), Web GUI on127.0.0.1:3080modlens-opencode-go/ modeldeepseek-v4-flash,reasoningEffort: max(a pure-text model; the provider declaresreasoning_contentas the thinking signature)request/headerrecords are identical across turns in both sessions)3. Symptom (user-visible)
In affected turns, a long block of English reasoning monologue appears as ordinary transcript text (rendered as markdown), instead of being collapsed under the "Think" disclosure. In one affected turn the leaked thinking was ~2,943 of a 3,027-character text block — i.e., the visible reply occupied only the last ~84 characters. In another observed case the reply itself was truncated at the per-turn output cap, consistent with the leaked thinking inflating the output.
4. Evidence
4.1 Session A — the first turn right after the rc.6→rc.7 GUI restart (2026-08-17, session log
session-1cf6231b-…/session.jsonl.zstd)Chunk-channel event counts per turn (parsed from the session record):
reasoning-chunksand 73text-chunks— the reasoning stream was folded into the text channel.assistant/messagefor turn 4 has noreasoningblock in any step; content shape is[text, tool-call](e.g. seq 17259, 17820, 17946). The step-1textblock is 3,027 characters: ~2,943 characters of English thinking + the ~84-character visible reply at the tail.replayState.blocksfor those messages is[text, tool-call]— the reasoning entry is missing there too (the same turn in the same session's earlier steps, e.g. seq 1733, showsblocks: [reasoning, text, tool-call]).4.2 Session B — days later, stable rc.7 runtime, two agentic turns (2026-08-18, session log
session-2cf17b41-…/session.jsonl.zstd)Chunk-channel event counts for the affected neighborhood:
Nine
assistant/messageentries in those two turns have content shape[text, tool-call, …]where thetextblock is the step's reasoning monologue and the step's actual visible reply is absent:dsh/client.jsline 799-800 …"Full structure of one leaked message (seq 487889, turn 159 step 2):
These turns are hours after the restart and are ordinary multi-tool-call agentic turns — disproving the earlier "one-time handover transient" hypothesis.
5. Analysis
5.1 What is ruled out
dsh-client-ui-conversation/dsh-client-ui-primitivesbundles are byte-identical inReasoningRow,DisclosureRow, and thecase "reasoning" → <ReasoningRow>dispatch; the UI faithfully renders what it receives. A message without areasoningblock cannot show a folded Think row.request/headerrecords for healthy and leaking turns are identical (provider: modlens-opencode-go,model: deepseek-v4-flash,reasoningEffort: max). Nothing changed between the healthy turn 5 (session A) and the leaking turn 4, nor between turn 158/160 and turn 159/161 (session B).5.2 What is consistent with a classification bug in the assembly layer
The same signature appears at three layers for the same turn, all pointing at where streaming deltas are classified into block types:
reasoning-chunks= 0 (or near-0) whiletext-chunksexplodes.assistant/messagecontent has noreasoningblock; the text block contains the thinking.replayState.blocks=[text, tool-call]— the reasoning entry is dropped there as well.If the mis-tagging happened at render time, layers 1–3 would still contain the reasoning. They do not. The classification is lost before storage — i.e., in the BlockAssembler / adapter seam that maps the provider stream (including
reasoning_content) to typed blocks.5.3 Trigger pattern (inference, n=2 sessions)
All affected turns share two properties:
tool-call-chunksin the leaking turns of session B; healthy turns have 0–23.Session A's turn 4 also had tool calls in every step. The working hypothesis: when reasoning deltas and tool-call deltas interleave heavily in one turn, the seam that decides "this delta is reasoning vs text" loses the reasoning classification for spans around the interleaving, and those spans fall into the text block. The 10 surviving reasoning-chunks in turn 159 vs 0 in turn 161 suggest it is a boundary-dependent behavior, not a hard failure.
5.4 Relation to the replay refactor
Session A's turn 4 was the first model call of the new process — the exact point where the rc.7 ReplayEnvelope change (#2596, flat v1
replayStatebeing taken over by the new envelope format) executes its handover. The v1→v2 transition is one candidate place where the reasoning entry of the envelope can be dropped while the text entry survives. This is consistent with the observedreplayState.blocks = [text, tool-call].6. Impact
7. Expected behavior
Reasoning streamed by the provider (
reasoning_content) should be classified intoreasoningblocks and folded under "Think" — for every turn shape, including multi-tool-call turns — matching healthy turns.8. Suggested check points for maintainers
packages/llm/llm/src/assembler.ts) delta classification when a turn interleaves reasoning and tool-call streams; specifically how a reasoning span adjacent to tool-call boundaries is typed.modlens-opencode-go-style routes (provider declaresreasoning_content).9. How to reproduce / verify locally
node decode-session.mjs <session.jsonl.zstd>(fzstd).reasoning-chunks≈ 0 with inflatedtext-chunks.assistant/messageentries: leaking steps havecontent = [text, tool-call]with thetextblock containing English planning prose.replayState.blocks:[text, tool-call]instead of[reasoning, text, tool-call].10. Notes
All reactions