deepseek-v4-pro emits Anthropic-style <invoke> tool-call text into content, not normalized to tool_calls — leaks to the user
#2158
Replies: 2 comments
|
Measured the "what happens today" half of this, so the feature discussion can start from ground truth: Fault injection into the real pi-ai accumulator (unmodified client/parser; only wire bytes synthetic): a completion whose This is one member of a family we've been mapping (model output lands in the wrong channel): markdown-fenced The prompt-side mitigation (system-prompt nudge to always use native tool_calls) reduces frequency but is not a guarantee — the measurement above is what happens on every miss. |
这个「输出风格泄漏」问题(来自 dsh-translate 与 dsh-defend 的实践):
如果是大模型问题,欢迎贴出模型与具体输出样例,帮你判断「到底要不要重打包 tool_calls」。 |
Uh oh!
There was an error while loading. Please reload this page.
Summary
When the agent runs on the
deepseek-officialroute with a reasoning model (deepseek-v4-pro, thinking enabled), the model sometimes expresses a tool call as plain text in the message content using Anthropic Claude's tool-use syntax, instead of the nativetool_callsfield:dsh-agent-looponly dispatches nativetool_callsblocks (seeexecuteToolCalls), and nothing in the loop / tool-presentation / LLM adapter strips or normalizes this text, so it ends up visible to the user as raw markup in the assistant's reply.Environment
dsh0.1.0-rc.6 (all@deepseek-ai/dsh-*packages at 0.1.0-rc.6)@deepseek-ai/dsh-llm-deepseek— provider routedeepseek-officialdeepseek-v4-pro(reasoning;thinkingdefaultenabled)@deepseek-ai/dsh-mcp-client— tools exposed asmcp__<serverName>__<rawName>Steps to reproduce
dshwithdsh-llm-deepseek(deepseek-official, deepseek-v4-pro, thinking enabled) and one or more MCP clients.toolCallTimeoutMs), then ask the agent to retry/continue.Observed: the model falls back to emitting the retry tool call as literal
<invoke>…</invoke>text in its content, which is rendered to the user verbatim.Expected vs actual
<invoke>text intotool_callsand execute it, or (b) strip it from the user-facing output so it never leaks.<invoke name="mcp__stock__wait">…</invoke>block appears in the assistant reply, and the tool is not executed.Root cause (from reading the source)
@deepseek-ai/dsh-agent-loopdispatches only nativetool_callsblocks (executeToolCallsinlib/index.js); there is no text-level tool-call parser.dsh-agent-loop,dsh-agent-tool-presentation,dsh-llmanddsh-llm-deepseekfinds no handling of the string<invoke.@deepseek-ai/dsh-mcp-clientconfirms the model-facing namemcp__<serverName>__<rawName>(publicToolName), which is exactly what appears insidename="…".So this is a gap: reasoning-model text-form tool calls (Claude-style
<invoke>) are neither normalized nor sanitized by the harness.Context / prior art
This is a known behavior of DeepSeek reasoning models — under ReAct-style prompting they can emit Anthropic tool-use syntax in text. In our own in-house ReAct loop we already handle it by stripping
<invoke>…</invoke>in a post-processing stage, which is why we recognized it. It would be great if the harness handled this natively.Suggestion
Either:
dsh-agent-loopthat parses<invoke name="…">…</invoke>(and<tool_calls>/function-call variants) intotool_callsblocks before dispatch; orHappy to provide logs or a minimal repro repo if a template exists.
All reactions