[bug]「只有 reasoning、无可见正文也无工具调用」的响应被判为成功 ⇒ 静默空回复 #7123
Replies: 3 comments
|
Thanks for the careful report — the classification table (推理退化 vs 流中断) and the 1. This shape already has a shipped plugin fix (it is the same gap as #6218)The first report of exactly this shape was discussion #6218 (reasoning-only
I installed that artifact from the registry into a clean project next to
Also checked in the same run: the reasoning bytes pass through untouched (the rewrite replaces the terminal chunk only, stream length unchanged, nothing is buffered), a healthy attempt (reasoning + visible text) is byte-identical, a tool-call attempt is untouched, and the shape-2 half from #6948 still fires. One structural point that saves you the second patch: the plugin sits at the 2. Correction: suggestion 1 would not change anythingI measured the two helpers you named (
Second, independent problem: both 3. Suggestion 2 is already satisfied by the retry pathThe retry decision happens before the assistant message is assembled, so a corrected degenerate completion is never written back into the conversation:
What a retried attempt does persist is (Where your hypothesis would be observable: the plugin's Two boundaries worth knowing
If you would rather not mount a plugin, the upstream fix you propose still stands and is the better long-term shape; the plugin is the version that works today. Its 0.2.1 peer range claims the 0.1.3-alpha.2, 0.1.5.x and 0.1.6.x lines, and I re-ran the two probes above against one release per line ( |
|
是有意的,而且被测试锁住了:
为什么这个默认值说得通关键在"干净的 EOF"这四个字。同一个用例喂的是
对比之下,你主报的那类空回复能进白名单,是因为它的判定落在终止块上、结论是"这次没有产出可见内容",重试方向明确; 想让它自动重试改 provider 路由上的 retryPolicy:
mode: normal
retryableCodes: [EMPTY_RESPONSE, RATE_LIMIT, SERVER, TIMEOUT, TRANSPORT, STREAM_CLOSED]两点要注意:
|
|
补充本机量化证据(与本帖同族:接受侧把"无实质正文"的结束当成功): 环境:macOS 自研 WKWebView 桌面壳 + dsh web(:3080),0.1.6-alpha.2 系 fork。
建议:成功/停止守卫按"实质答案"判定而非"是否有可见文本":末条为旁白体或纯 reasoning 且本回合存在未完成工具意图时,判 incomplete 并自动续跑一次或提示重试;否则长任务中模型提前 stop 会把过程旁白留作"最终答案"。 |
Uh oh!
There was an error while loading. Please reload this page.
[Bug] 只含 reasoning 的响应被判为成功完成 ⇒ 静默空回复 / Response with only reasoning blocks is reported as completed ⇒ silent empty reply
摘要
当模型产出一条只含
reasoningblock、既没有text也没有tool-call的最终消息时,Harness 把它当作正常完成:turn/end记completed、不报错、不重试、日志里没有任何异常痕迹。用户侧表现为「思考区滚完,正文一个字都没有」——看起来像客户端卡死或断流,而实际请求被判为"成功"。同一判据缺陷在两个适配器里各出现一次(见「同一缺陷的第二个适配器」)。
环境
0.1.5-rc.2,Windows 11deepseek-official,modeldeepseek-v4-flashreasoningEffort: high,maxTokens: 256000证据
退化消息的共同铁证是
outputTokens恰好等于reasoningTokens⇒ 整个输出预算都花在 reasoning 上,可见正文零 token:推理内容本身呈自我催促循环(复读),例如结尾形如:
与「流中断」的区别(同一现象,两个成因)
审计同一批日志时还发现另外 2 条「只有 reasoning」的消息,但它们是另一种成因,混为一谈会开错处方:
usage,且outputTokens === reasoningTokens;推理文本呈低熵复读data.interrupted === true,或整条没有usage;推理文本正常,只是被硬切断(本次两条的思考停在半条 shell 命令上)相关因素(统计,非因果断言)
扫描 18 个会话日志、共 2584 条
assistant/message:根因定位(源码)
1.
@deepseek-ai/dsh-llm-deepseeklib/index.js的translate()处理[DONE]哨兵处:判据是
order.length === 0(一个 block 都没打开)。而 reasoning block 会先open("reasoning")并进入order,因此**「只有 reasoning」的消息order.length === 1** ⇒ 直接走reason(stop)⇒ 被判为成功。另外,
@deepseek-ai/dsh-llm已导出assistantStreamHasVisibleContent/assistantStreamHasVisibleText(见该包lib/index.js导出表),但本 adapter 未使用它们。2.
@deepseek-ai/dsh-llm-pi-ai(同一缺陷的第二个适配器)lib/index.js的mapStopReason():同样是只看「块数为 0」。而该 adapter 的
content里有思考块(pi-ai 原生 message 用type: "thinking";completeTerminalMessage()从流块恢复时用 DSH 风格的type: "reasoning")⇒ 只要模型思考过,content.length ≥ 1⇒ 判stop⇒ 同样静默成功。修判据时需注意:这里两种类型名都要认(工具块为
toolCall或tool-call,正文块两条路径都是text)。影响
turn/end仍是completed),无法自查;重发也常继续退化(见上表连续五轮);[DONE])会抛STREAM_CLOSED,那条路径至少有错误。建议
dsh-llm-deepseek可直接改用现成的assistantStreamHasVisibleContent),并映射为EMPTY_RESPONSE—— 该错误码已在默认重试白名单内(@deepseek-ai/dsh-llm/lib/types/retry-policy.js的DEFAULT_RETRYABLE_CODES),于是这类空回复能被自动重试救回,而不是静默失败。guard/家族(timeout-policy、repeat-tool-reminder)都以工具调用为中心,而退化的典型形态恰恰是不发工具调用。若在 agent-loop 层增加一个「连续 N 步无可见产出」的守卫,可在重试之外更早地打断(这一层的事前干预比事后重试更省 token)。STREAM_CLOSED(真正的流截断)不在DEFAULT_RETRYABLE_CODES里,因此真断流不会自动重试(对比:TRANSPORT在内,故传输类失败会自动重试)。是否有意如此?若纳入,用户侧"说一半没了"的问题也能显著减少。自查方法
扫描会话日志(多帧 zstd,需逐帧解压)中所有
assistant/message记录,判定:配合
inputTokens + cacheReadTokens(上下文体积)即可定位高发区间。2026-09-18-dsh-reasoning-only-empty-message.md
All reactions