Replies: 4 comments
Tool-scheduler crash on the first tool call under
|
| launch | import.meta.resolve('@deepseek-ai/dsh-tools') |
bare.SYM === lib.SYM |
|---|---|---|
node --import tsx/esm … |
packages/core/tools/src/index.ts |
false |
plain node (built entry) |
packages/core/tools/lib/index.js |
true |
Why the error message does not lead anywhere. Both symbols have the same .description (@deepseek-ai/dsh-tools.scheduler), so nothing in the output distinguishes them; agent.ts:334 flattens any non-LlmError into {message, code:'UNKNOWN'}, dropping the stack; and the message names prepare, which invites you to hunt for a missing method rather than a mismatched key.
Why earlier calls appear to work. ctx.tools.executionMode(...) is a prototype method, so it resolves through the prototype chain and is unaffected by the symbol. Only the symbol-keyed own field is missing. tool/call events are appended normally right up to the failing prepare.
Root cause 2 — the orphan is never repaired
The crash sits between two writes:
callSeqs[index] = appendToolCall(session, turn, step, call.block) // durable
started++
const prepared = await ctx.tools[TOOL_RUNTIME_SCHEDULER].prepare(call.exec) // throws herepackages/core/session/src/repair.ts is built to close exactly this kind of gap: interruptedTurnClosers() emits synthetic error results for calls with no recorded outcome. But it returns [] immediately when openTurn === null — it only handles unclosed turns. A scheduler crash closes the turn with {kind:'error'}, so repair skips it and the unpaired call stays in the log forever.
From then on, packages/llm/llm-deepseek/src/protocols/messages/serialize.ts:117 rejects every request:
if (pending.size > 0) throw new LlmError('DeepSeek Messages tool calls need immediate results', 'INVALID_REQUEST')This wedge is Messages-protocol-only. The chat-completions serializer does not enforce call/result pairing, so the same log keeps working when llm-deepseek.protocol is chat-completions. Worth knowing when building a repro.
Why the existing test/CI matrix does not catch this
The two halves of the divergence are each covered, but never together — the hybrid state is not exercised anywhere:
.github/workflows/e2e.yml:109-123builds and then runs e2e withDSH_EXAMPLE_MODE: lib— a purelibgraph.- The coverage lane in
ci.ymldeliberately builds nothing, so workspace imports resolve tosrc— a puresrcgraph. apps/cli/tests/source-launch.compat.spec.tsdoes launch the real CLI through tsx, but it only asserts the missing---profilediagnostic and exits before any tool call or plugin dynamic import.
docs/testing.md:45 already names this exact hazard — "stale artifacts there load a second copy of module singletons" — but scopes the rule to the vitest process and explicitly carves out subprocesses ("Built artifacts are consumed only explicitly: lib-mode subprocesses…"). The Loader's dynamic import inside a src-mode child process is precisely the unguarded case.
Nothing checks lib/ against src/; there is no host/CLI watch or rebuild loop; and apps/cli/README.md:56 states plainly that "The launcher does not check freshness."
Finally, the type system cannot help here: ctx.tools[TOOL_RUNTIME_SCHEDULER] is declared non-optional, which is correct under a single module graph. Types are a compile-time claim about one module graph and cannot see that the module was evaluated twice — so no call site is written defensively and no check reports the possibility.
Suggested fix
1. Make the handshake survive duplicate module evaluation (one line). packages/core/tools/src/index.ts:463:
-export const TOOL_RUNTIME_SCHEDULER: unique symbol = Symbol('@deepseek-ai/dsh-tools.scheduler')
+export const TOOL_RUNTIME_SCHEDULER: unique symbol = Symbol.for('@deepseek-ai/dsh-tools.scheduler')Symbol.for uses the global registry, so both module instances hand out the same key. Verified: the identity comparison above goes from false to true under tsx, and the crash no longer reproduces.
2. Do not leave an announced call without a result. In tool-calls.ts, the scheduler-failure path rethrows without closing out started calls (the abort path already does this via appendSkippedToolCall). Closing every remaining call of the step before propagating keeps the transcript serializable. This is defense in depth — it does not fix (1), but it stops any internal failure from poisoning a session beyond the one turn.
3. Optional: assert at activation that ctx.tools?.[TOOL_RUNTIME_SCHEDULER] !== undefined, so a future instance split fails loudly at boot instead of at the session's first tool call.
Verification performed
- Symbol identity probe in both launch modes (table above).
- A repro that boots the real CLI in tsx mode against a mock LLM and issues a tool call: crashed before the fix, runs two tool calls cleanly after.
- For defect 2, since
packages/test-support/llm-mock-serverserves chat-completions only (it 404s/v1/messages), the A/B was done offline at the serializer boundary using the real readers: decode the session log, cut each variant at the request boundary (the firstuser/messageafter the failedturn/end), rebuild withSession.create(id, events), project withderiveMessages(), then call the Messagesserialize(). On the identical injected failure: without the guard it throwstool calls need immediate results; with the guard it serializes to 3 wire messages. packages/core/agent-loop+packages/core/tools(842 tests) andpackages/core/session+packages/llm/llm-deepseek+packages/session(3026 tests) all pass with both changes. No test encoded the previous "do not fabricate tool results" behaviour.
中文版
概要
两个缺陷叠加。第一个是一行修复;第二个把它从"一次崩溃"升级成"会话日志的永久性损坏"。
-
崩溃。 在构建产物已存在的情况下,用 tsx 源码模式启动 CLI(
pnpm dsh …)会让@deepseek-ai/dsh-tools被求值两次。TOOL_RUNTIME_SCHEDULER用的是普通Symbol(),两份模块实例于是持有不同的键,ctx.tools[TOOL_RUNTIME_SCHEDULER]为undefined,会话的第一次工具调用即崩溃:Cannot read properties of undefined (reading 'prepare')。 -
永久卡死。 崩溃点落在
tool/call事件已持久化之后、任何tool/result之前。repair.ts只修复尚未闭合的 turn,而调度器崩溃会把 turn 以{kind:'error'}正常闭合,因此这条孤立调用永远不会被修复。之后每一次请求都会失败于DeepSeek Messages tool calls need immediate results——永久如此,重启服务器也没用,因为损坏在磁盘上的会话日志里。
根 README 的安装流程(pnpm install → pnpm run build → pnpm dsh web)产生的正是触发第 1 条的状态。
环境
dshv0.1.6-alpha.2,commitddefc45- Node
v24.17.0,pnpm11.7.0 - Windows 11(缺陷在于模块解析,与平台无关)
- 启动方式:
pnpm dsh web(即node --import tsx/esm apps/cli/src/bin.ts web)
复现步骤
git clone https://github.com/deepseek-ai/deepseek-harness
cd deepseek-harness
pnpm install
pnpm run build # 关键:这一步让 lib/ 存在
pnpm dsh web # tsx 源码模式在任意会话中发一轮会调用工具的消息。预期:工具正常执行。实际:该轮失败于 Cannot read properties of undefined (reading 'prepare'),code 为 UNKNOWN;此后该会话的每一轮都失败于 DeepSeek Messages tool calls need immediate results。
不调用工具的轮次是正常的——这也是为什么第一轮常常看起来没问题。
根因 1 —— 两份模块实例,两个 Symbol()
packages/core/tools/src/index.ts:463:
export const TOOL_RUNTIME_SCHEDULER: unique symbol = Symbol('@deepseek-ai/dsh-tools.scheduler')packages/core/agent-loop/src/tool-calls.ts:170:
const prepared = await ctx.tools[TOOL_RUNTIME_SCHEDULER].prepare(call.exec)在 tsx 下同一个包名有两条解析路径:
- tsx 自己实现了一套 tsconfig
paths解析(堆栈里可见resolveTsPathsSync),把裸@deepseek-ai/*映射到src/。这是刻意的——正因如此源码树才能免构建直接运行。 - cordis 插件 Loader 通过自己带 base URL 的动态 import 挂载插件行(
packages/boot/app-boot/src/index.ts:510-528),绕开了上述映射,落到 packageexports→lib/。
两次求值,两个不同的 symbol。ToolRuntime 实例身上挂着 lib 那份 symbol 作自有属性键,而 agent-loop 手里是 src 那份,于是查表落空。
两个方向的实测:
| 启动方式 | import.meta.resolve('@deepseek-ai/dsh-tools') |
bare.SYM === lib.SYM |
|---|---|---|
node --import tsx/esm … |
packages/core/tools/src/index.ts |
false |
原生 node(构建入口) |
packages/core/tools/lib/index.js |
true |
为什么这条报错毫无指向性。 两个 symbol 的 .description 完全相同(@deepseek-ai/dsh-tools.scheduler),输出里没有任何东西能区分它们;agent.ts:334 把所有非 LlmError 压成 {message, code:'UNKNOWN'},堆栈丢失;而报错文案点名的是 prepare,会诱导你去查"哪个方法没了",而真正原因是"键对不上"。
为什么之前的调用看起来是好的。 ctx.tools.executionMode(...) 是原型方法,走原型链解析,不受 symbol 影响。丢的只是 symbol 键的自有字段。tool/call 事件会一直正常写入,直到那个失败的 prepare。
根因 2 —— 孤立调用永远不会被修复
崩溃点卡在两次写入之间:
callSeqs[index] = appendToolCall(session, turn, step, call.block) // 已持久化
started++
const prepared = await ctx.tools[TOOL_RUNTIME_SCHEDULER].prepare(call.exec) // 在这里抛出packages/core/session/src/repair.ts 本就是为补这类缺口而生的:interruptedTurnClosers() 会为"没有记录结果的调用"补发合成 error result。但它在 openTurn === null 时直接返回 []——它只处理未闭合的 turn。而调度器崩溃会把 turn 以 {kind:'error'} 正常闭合,于是 repair 跳过,未配对的调用永久留在日志里。
此后 packages/llm/llm-deepseek/src/protocols/messages/serialize.ts:117 会拒绝每一个请求:
if (pending.size > 0) throw new LlmError('DeepSeek Messages tool calls need immediate results', 'INVALID_REQUEST')这个卡死只在 Messages 协议下发生。 chat-completions 的序列化器不校验调用/结果配对,所以同一条日志在 llm-deepseek.protocol 为 chat-completions 时依然能跑。构造复现时值得注意这一点。
为什么现有的测试/CI 矩阵抓不到
分歧的两半各自都有覆盖,但从不一起覆盖——这个混合状态在任何地方都没有被执行到:
.github/workflows/e2e.yml:109-123先构建,再用DSH_EXAMPLE_MODE: lib跑 e2e —— 纯lib图。ci.yml的覆盖率 lane 刻意不构建,于是工作区导入全部解析到src—— 纯src图。apps/cli/tests/source-launch.compat.spec.ts确实通过 tsx 启动了真实 CLI,但它只断言缺少--profile时的诊断信息就退出了,走不到任何工具调用或插件动态 import。
docs/testing.md:45 其实已经点名过这个危险——"stale artifacts there load a second copy of module singletons"——但该规则被限定在 vitest 进程内,并明确把子进程排除在外("Built artifacts are consumed only explicitly: lib-mode subprocesses…")。而 Loader 在一个 src 模式子进程内部做的那次动态 import,恰恰是这个规则没覆盖到的情况。
没有任何检查比对 lib/ 与 src/;没有 host/CLI 的 watch 或 rebuild 循环;apps/cli/README.md:56 也直说了:"The launcher does not check freshness."
最后,类型系统在这里帮不上忙:ctx.tools[TOOL_RUNTIME_SCHEDULER] 被声明为非可选,这在单模块图下是完全正确的。类型是关于一份模块图的编译期断言,它看不见模块被求值了两次——于是没有任何调用点写防御性判空,也没有任何检查会报告这种可能性。
建议的修复
1. 让握手在模块被重复求值时依然成立(一行)。 packages/core/tools/src/index.ts:463:
-export const TOOL_RUNTIME_SCHEDULER: unique symbol = Symbol('@deepseek-ai/dsh-tools.scheduler')
+export const TOOL_RUNTIME_SCHEDULER: unique symbol = Symbol.for('@deepseek-ai/dsh-tools.scheduler')Symbol.for 走全局注册表,两份模块实例会拿到同一个键。实测:tsx 下上面的同一性比较由 false 变为 true,崩溃不再复现。
2. 不要让一条已公布的调用没有结果。 在 tool-calls.ts 中,调度器失败路径在重抛前没有收尾已启动的调用(而 abort 路径已经通过 appendSkippedToolCall 做了这件事)。在向上抛之前把该 step 剩余的全部调用收尾,可以保证记录始终可序列化。这属于纵深防御——它修不了第 1 条,但能阻止任何内部失败把会话毒化到超过一轮。
3. 可选: 在激活时断言 ctx.tools?.[TOOL_RUNTIME_SCHEDULER] !== undefined,让将来任何实例分裂在启动时就大声失败,而不是等到会话的第一次工具调用。
已做的验证
- 两种启动模式下的 symbol 同一性探针(见上表)。
- 一个在 tsx 模式下启动真实 CLI、对接 mock LLM 并发起工具调用的复现:修复前必崩,修复后两次工具调用全部正常。
- 针对缺陷 2:由于
packages/test-support/llm-mock-server只提供 chat-completions(对/v1/messages返回 404),A/B 改为在序列化器边界离线进行,且全部使用真实代码路径——解码会话日志、把两个变体都截在请求边界上(失败的turn/end之后的第一条user/message)、用Session.create(id, events)重建、用deriveMessages()投影,再调用 Messages 的serialize()。在完全相同的注入失败下:无护栏版本抛出tool calls need immediate results,有护栏版本正常序列化为 3 条 wire message。 - 两处改动后,
packages/core/agent-loop+packages/core/tools(842 项)与packages/core/session+packages/llm/llm-deepseek+packages/session(3026 项)测试全绿。没有测试固化了此前"不伪造工具结果"的行为。
|
原来不止我一个遇到这个问题,发版都不测的吗... |
|
Confirming the diagnosis from a second machine, plus two additions that matter for whoever applies the fix. What I reproducedI worked from a checkout at 1. The declaration is still plain — so this is not fixed yet.
export const TOOL_RUNTIME_SCHEDULER: unique symbol = Symbol('@deepseek-ai/dsh-tools.scheduler')2. Under tsx the bare specifier resolves to source, and a second evaluation produces a different symbol — same process, measured: 3. The hazard is not really about src-vs-lib. Two instances of the same built file reproduce it, which is the cleaner demonstration and also confirms the remedy: Addition 1 — do not blanket-replace
|
|
This thread has the mechanism and the remedy right. Three things are still missing from it: why 0.1.5-rc.2 → 0.1.6-alpha.2 is a boundary at all (the README sequence does not explain it — 1. The boundary is a one-commit default flip — and that flip implemented written intent
apps/cli/src/profile-boot.ts:263
- const resolutionMode = packaged ? 'runtime' : options.resolutionMode ?? 'link'
+ const resolutionMode = packaged ? 'runtime' : options.resolutionMode ?? 'runtime'
apps/desktop-host/src/index.ts:21
- resolutionMode: process.argv[5] === 'runtime' ? 'runtime' : 'link',
+ resolutionMode: 'runtime',That flip is the whole release delta, not the build step: With The same commit's other half was reverted 39 minutes later. But this is not a stray edit, and that changes what the fix should be. The three layers are not alternatives:
(b) is the one that generalizes, and it is this repo's own convention rather than a workaround: 2. Defect 2 — the closer set is
|
Uh oh!
There was an error while loading. Please reload this page.
Environment
Summary
Tool calls crash in 0.1.6-alpha.2. Simple conversation works fine, but the moment the model tries to invoke any tool (shell, computer-control, etc.) the turn fails. 0.1.5-rc.2 works correctly on the same machine.
Steps to reproduce
Expected
The tool call executes and returns "hello".
Actual
The
--jsonevent stream shows the model emits atool_call, then the turn ends with:{"type":"status","phase":"turn_end","reason":{"kind":"error","error":{"message":"Cannot read properties of undefined (reading 'prepare')","code":"UNKNOWN"}}}Root cause (source analysis)
packages/core/agent-loop/src/tool-calls.ts:170:At runtime
ctx.tools[TOOL_RUNTIME_SCHEDULER]isundefined— the Symbol-keyed scheduler view (TOOL_RUNTIME_SCHEDULER, defined inpackages/core/tools/src/index.ts) is not reachable, so.preparethrows.Notes
pnpm run build, client build record present).All reactions