[Bug] workflow 工具调用永不返回,父 Agent 卡死在 step/start 后无事件 #8827
Replies: 7 comments
你这条的形态很清楚:"嵌套等待"——父 Agent 的取消/超时没有覆盖子阶段的推进1. 从你的描述能确定的三件事你写: ⇒ 三条合起来说明不是"工具崩了",而是"工具的返回信号丢了":
这类缺陷最典型的成因有两个方向,建议你在报告里明确区分:
判定方法很便宜:看第 5 阶段是谁启动的——若是父在等待期间另外派生的,那更可能是 A(完成的定义与实际阶段数不一致)。 2. 建议的诉求(两条,第二条是根治)
"无事件"这一点值得单列:你说 38 分钟一个新事件都没有——这意味着不仅工具没返回,连"还在等"都没被表达。这比"慢"严重得多。 3. 关于可复现性(这条需要你补)你现在给的是一次现场。⇒ 请补:
4. 版本口径你写 DSH Desktop 一条边界我给出的是按你描述推出的"嵌套等待"定性与两个候选方向(A 计数 / B 通道)。具体是哪一个要看第 3 节那三样——我没有你的 workflow 定义,也没有该会话的事件流。 |
|
补充数据(原发帖人):你要的三项,以及两处必须更正的事实 先把结论放前面:用你给的判据(阶段数 / 第 5 阶段来源 / 是否稳定重跑)把 journal 逐条对回去之后,原报告里"丢失完成信号、父 Agent 永久卡死"的判断不成立。实际情况是:这个 workflow 声明了 5 个阶段(不是 4 个),第 5 阶段在被重启掐断时仍在正常推进;从调用发起到最后一条记录共 37 分 48 秒,其间父会话一直有阶段事件。所以你的候选 A(计数差一)和候选 B(完成事件发了没收到)都没有事实基础。 但这条报告并非无效:真正可确认的是另外三条缺陷(无超时 / 阶段内无进度信号 / 无法中断),我按代码逐条给了位置。下面全部是原始记录,可自行复核。 1. 你要的 workflow 定义调用标识:
{
"description": "Deep source recon of DeepSeek web reverse-engineering projects",
"name": "ds-web-recon",
"phases": [
{ "title": "ds2api", "detail": "Go source + SSE spec docs" },
{ "title": "snake", "detail": "expert mode / pow / files" },
{ "title": "wu-jiyan", "detail": "pow_solver + adapter" },
{ "title": "web2api-others", "detail": "extra projects and endpoint catalogs" },
{ "title": "risk", "detail": "breakages, dates, maintenance risk" }
]
}
const COMMON = `…`; // ~2.6 KB 通用 recon 指令
phase("ds2api"); const ds2api = await agent(`…`, { label: "ds2api source recon", phase: "ds2api" });
phase("snake"); const snake = await agent(`…`, { label: "snake expert-mode recon", phase: "snake" });
phase("wu-jiyan"); const wu = await agent(`…`, { label: "wu-jiyan + Fly143 recon", phase: "wu-jiyan" });
phase("web2api-others"); const others = await agent(`…`, { label: "cross-check other projects", phase: "web2api-others" });
phase("risk"); const risk = await agent(`…`, { label: "maintenance risk + blockers", phase: "risk" });
return { ds2api, snake, wu, others, risk };⇒ 5 个阶段、纯串行、逐个 对照:同会话更早的 2. 第 5 阶段的来源(你的问题 b)同一 seq=5 紧跟在 seq=4 完成之后 44 毫秒、由 workflow 自己按 3. 完整时间线(UTC,毫秒级)
4. 两处更正(1) "此后 38 分钟无任何新事件" → 38 分钟是调用总时长,不是静默期。 原帖真正站得住的表述是:缺 step 级事件——无 (2) "永久卡死" → 在最后一条记录的时刻,运行仍在进行。 即:它不是"完成事件丢了",而是第 5 阶段根本没结束,进程先被销毁了。按 4 个已完成阶段的耗时分布(7–10 分钟),第 5 阶段当时才跑了 3 分 27 秒,完全在正常区间内。 ⇒ 结论:这不是"卡死",是"一个合法的长任务被观察者判成卡死并强杀"。原帖的三信号判据( 5. 代码级:为什么它"永远等得下去"(这部分是确认,非推断)
6. 新版核对:0.2.1-alpha.1 上未修复
顺带更正一处环境口径: 7. 可复现性:目前给不出"稳定复现"
8. 建议把诉求改写成这四条
附:本次取证方式(可复现)会话日志是多帧 zstd(本例父会话 52 帧),按 magic 最后一句:我认可你对"等待必须可超时/可取消"的判断,而且这次样本正好证明了它——只是原因不是"完成信号丢了",而是"它还在干活,没人告诉观察者这件事,也没有任何机制叫停"。 |
|
追加更正(同一位发帖人):上一条把"进程被销毁"讲得太粗——这次不是崩溃,是被中断取消的。 我在父会话里找到了这次 同时有一条用户消息「逐步暂停」在 20:05:33 入队,并在 20:08:15 被取消。⇒ 推断(强,但日志只给了 这条数据把上一轮的两点补完整:
顺便补两个可复核的数据点(回答第 3 节的可复现性):
|
Supplementary data (original reporter) — English versionThis is the English version of my two earlier comments (data + one correction to my own follow-up). Everything below is reproducible from the session journals on the reporting machine. 1. Headline: the "permanent hang" reading does not survive the raw journalThe workflow in question declares 5 phases, not 4, and its 5th stage was still running normally when everything was interrupted. Both of the reviewer's candidate causes are ruled out by the data:
The real, code-backed defects are different (see §5): no timeout, no in-stage progress signal, and a cancel path that did not durably write a result. 2. The workflow definitionCall:
{
"description": "Deep source recon of DeepSeek web reverse-engineering projects",
"name": "ds-web-recon",
"phases": [
{ "title": "ds2api", "detail": "Go source + SSE spec docs" },
{ "title": "snake", "detail": "expert mode / pow / files" },
{ "title": "wu-jiyan", "detail": "pow_solver + adapter" },
{ "title": "web2api-others", "detail": "extra projects and endpoint catalogs" },
{ "title": "risk", "detail": "breakages, dates, maintenance risk" }
]
}Control flow of const COMMON = `…`; // ~2.6 KB shared instructions
phase("ds2api"); const ds2api = await agent(`…`, { label: "ds2api source recon", phase: "ds2api" });
phase("snake"); const snake = await agent(`…`, { label: "snake expert-mode recon", phase: "snake" });
phase("wu-jiyan"); const wu = await agent(`…`, { label: "wu-jiyan + Fly143 recon", phase: "wu-jiyan" });
phase("web2api-others"); const others = await agent(`…`, { label: "cross-check other projects", phase: "web2api-others" });
phase("risk"); const risk = await agent(`…`, { label: "maintenance risk + blockers", phase: "risk" });
return { ds2api, snake, wu, others, risk };Strictly serial, 5 3. Origin of the 5th stage (reviewer question b)The 5th 4. Timeline (UTC, ms) and per-stage durations
Two corrections to the original report(1) "No events for ~38 minutes" is false. 38 minutes is the total call duration (11:30:27.388 → 12:08:15.880 = 37 m 48.5 s), not a silent window. Inside it the parent session recorded 18 events: 1 × (2) The run was still in flight, and it ended by interruption — not by a hang. The 5th-stage child session And in the parent session of that subagent (the root session) the same moment reads: {"type":"tool/result","seq":879,"time":1791115695845,"data":{"turn":6,"step":5,"message":{"role":"tool","isError":true,"content":[{"type":"text","text":"The tool call was interrupted after it was recorded, but no result was durably recorded. Its outcome is unknown. …"}]}}}
{"type":"turn/end","seq":881,"time":1791115695845,"data":{"turn":6,"reason":{"kind":"interrupted"}}}A user message 5. Code: why the wait can never time out (confirmed)
6. Newest alpha: unchanged
Every line cited in §5 is byte-identical in 7. Reproducibility: no data pointAcross all 102 session journals on this machine there are 9 8. Suggested restatement of the asks
Method note (for anyone re-verifying)Session logs are multi-frame zstd (52 frames in this parent journal, 893 in another). |
附:workflow
|
我接受你的两处更正——上一轮我给出的两个候选方向都被你的原始记录否掉了1. 撤回你的数据把我的判断推翻了,而且推翻得干净:
⇒ 所以我上一轮那句"完成信号丢失、父 Agent 永久卡死"是错的,请以你的记录为准。我本该先要那三样再下判断,而不是先给两个候选方向——这是我的方法错误,记下了。 2. 但真正的根源浮出来了:取消没有传播到子代理把你的更正与另外两条报告并排看,指向的是同一个根因:
⇒ 根因判断:子代理的生命周期与取消信号,没有和父级 turn 绑定。于是"停"只能停在父级:父 turn 结束了,子代理还在跑;要真正停住只能重启进程。这解释了这三条报告的全部现象——它们不是三个独立缺陷。 3. 解决方案(三层,按可落地性)
4. 你这份补充本身的价值(建议这样定稿)你手上现在有三样别人没有的东西:完整
这样它就从一条被否掉的报告,变成一条有完整现场、有代码位置、可验收的缺陷报告。 5. 一条请求若方便,请把你核到的三处代码位置也贴进这条(我这边只看到你提到"按代码逐条给了位置")——有了它们,我可以对着当前 一条边界我确认的是你的记录与我的两个候选方向不相容(因此撤回),以及这三条报告在机制上同源这一推断。**"取消在哪一层断掉"**需要你补的代码位置才能定论——我没有该会话的 journal。 |
回复:你要的三处代码位置(完整版)+ 一处我建议不要采纳的新判断1. 缺陷 → 代码位置 → 验收判据(这就是你要的第 5 节)我上一条给的是散点,这里合并成一张表(位置为已安装的
第 1/2/3/4/5 条都是本次现场直接可证的;第 6 条是本次误判成"卡死"的直接原因;第 7 条是排除候选 A 的依据。 2. 关于「取消没有向下传播到子代理」——这次现场支撑不了,建议标为"未证"同一进程内的时间线(UTC): 子会话的最后事件早于父级
要把这条立住,需要的是一个父 turn 结束后进程仍存活、而子会话继续产生事件的现场——本次不是。建议把判据写成可证伪形式:
我认同这是值得查的方向(你并到 #8589 / #8849 的动机也成立),但请标注为"未证/待补",而不是"根因"。 3. 关于把报告重写成「取消不向下传播」不建议把根因替换掉,理由见第 2 节。建议的改写方向是:把标题与根因定在本次可证的两条上——
保留你已经认可的"原报告『永久卡死』判断不成立"作为正文首段;把"取消传播"作为并列的待查项(附上面那条可证伪判据)。这样报告就是"可验收"的,而不是把一条未证推断换成另一条。 Reply: the code locations you asked for, plus one claim I would not adopt1. Defect → location → acceptance criterion
Items 1–5 are directly provable from this incident; item 6 is the reason the run looked hung; item 7 is why candidate A was excluded. 2. "Cancellation did not propagate to the subagent" — this incident does not support itThe child's last event precedes the parent's A falsifiable criterion instead:
I agree this is worth investigating (and your grouping with #8589 / #8849 is a reasonable hypothesis), but please label it unproven/pending rather than root cause. 3. On rewriting the report around "cancellation does not propagate"I would not replace the root cause (see §2). A defensible rewrite anchors the title and cause on the two things this incident proves:
Keep your accepted retraction ("the original 'permanently hung' reading does not hold") as the opening paragraph, and add cancel propagation as a parallel open question with the falsifiable criterion above. That keeps the report acceptable-as-is instead of trading one unproven claim for another. |
Uh oh!
There was an error while loading. Please reload this page.
DSH 缺陷报告:
workflow工具调用永不返回,父 Agent 卡死在 step/start 后无任何事件摘要
一次
workflow工具调用在实际完成全部 4 个阶段后从未返回,导致父 Agent 永久卡死在step/start turn=1 step=8,此后 38 分钟无任何新事件产生。与此同时其内部启动的第 5 阶段子 Agent 仍在正常推进,形成嵌套等待:子等工具结果 → 父等子 → 父的调用永不返回。
环境
0.2.0-rc.2nightly(https://download.deepseek.com/dsh-desk/feeds/mac-arm64/)com.deepseek.dshcordisdeepseek-flash,agentReasoningEffort: "max"desktop症状(可观测判据)
出现以下三个信号同时成立且持续 >7 分钟时,即为该缺陷:
pendingCalls长期不回收session_projcache中sessionStats.pendingCalls含一个永不消失的callId。openStep: null且无新步骤Agent 既不在执行步骤,也未推进到下一步。
session.v4.jsonl.zstd字节数不变,对应.jsonprojcache 的 mtime 不再更新。对照:正常等待 LLM 长响应时,日志会持续增长(实测单次停顿上限约 1.5 分钟)。
复现过程(真实记录)
时间线
059dd7d9启动(subagent/descriptor,mode: continuable)workflow调用call_00_mveyI3BYkFz0SbKENggD7199(depth-probe)→ 正常返回workflow调用call_01_RoATNRhZVJMwcagNRMlh1276(ds-web-recon)发起tool-workflow/run-start name=ds-web-recontool-workflow/agent-start seq=1 'ds2api source recon'agent/inbox/spliced target=next-step(后台任务通知被塞入父 Agent 的 inbox)agent/inbox/spliced target=next-stepagent-end seq=1 completedagent-start seq=2 'snake expert-mode recon'agent-end seq=2 completedagent-start seq=3 'wu-jiyan + Fly143 recon'agent-end seq=3 completedagent-start seq=4 'cross-check other projects'agent-end seq=4 completedagent-start seq=5 'maintenance risk + blockers'agent/inbox/spliced target=next-step outcome=canceled removed=2关键异常点
①
workflow调用从未返回事件序列中:
seq 83tool/call name=workflow id=call_01_RoATNRhZVJMwcagNRMlh1276✅ 已发起seq 84tool-workflow/run-start name=ds-web-recon✅seq 89/92/95/98四个阶段agent-end ... completed✅tool-workflow/run-endtool/resultseq 101,没有step/end对比:第一次
workflow调用(depth-probe)有完整的run-start(69) → agent-start(71) → agent-end(72) → run-end(73) → tool/result(74)序列。② 父 Agent 卡在 step 8
sessionStats最终定格:steps: 7、openStep: null、pendingCalls: [call_01_RoATNRhZVJMwcagNRMlh1276]。③ 后台任务通知被投递到错误的 Agent
seq 87与seq 88的 inbox 通知内容是父 Agent 早期(step 6–7)通过tool/call name=bash启动的两个后台 job 的完成通知:这两个通知插入了父 Agent 的
next-stepinbox,而父 Agent 此时正阻塞在workflow调用上,无法消费它们。最终在seq 101被批量canceled(removed=2)。④ 子 Agent 在父卡死后仍正常推进
第 5 阶段子 Agent
4876e8fa的独立会话显示它一直在正常工作:即:父 Agent 早已卡死,子 Agent 毫无察觉地继续工作,只有外部重启才终止了整个级联。
影响
tool/result,父 Agent 永不复原。用户只能重启 App。父 Agent 却收不到任何产出。本例 7 阶段共产生约 2.1 MB 会话日志、耗时 38 分钟。
session.lock由 host 进程持有;重启后锁文件残留(0 字节,无进程持有)。
附带发现:子 Agent 没有任何终止手段
排查过程中确认以下几点,建议一并评估:
interrupt_agentError: active teammate "<id>" not foundlist_agentslead,不含任何 subagent,故无 target 可用@deepseek-ai/*包中无 CLI 入口)dsh-client-ui-sidebar/dsh-client-ui-session的lib/client.js、lib/index.js,无deleteSession/removeSession/closeSession/archiveSession等符号即:用户面对一个挂死的子 Agent,除了重启整个应用之外无计可施,且重启会中断所有工作区的会话。
期望行为(建议)
workflow调用必须有超时与失败返回无论内部阶段成功与否,
tool/call name=workflow都应最终产生tool/result(成功或
isError: true),并配套tool-workflow/run-end。单个 workflow 阶段超过阈值(建议可配置,默认如 10 分钟)应标记失败并继续/中止,
而非无限等待。本例单阶段最长 9 分 51 秒,串行 4 阶段共 34 分钟,仍在同一调用内。
当 Agent 正阻塞在某个 tool 调用上时,
next-stepinbox 的通知应排队等待该调用返回后再投递,而不是被投递后静默
canceled。至少支持
interrupt/kill单个 subagent 会话(工具层或 UI 层),避免只能靠重启 App 收场。
启动时清理
session.lock中未被任何进程持有的残留锁。在 UI 上暴露
pendingCalls+openStep状态,使"卡死"与"慢"可被用户区分(本例中两者外观完全相同,直到 38 分钟后才能确认)。
附:复现与取证方法
定位卡死会话
判断"真卡死" vs "正在等 LLM"
解码 zstd 多帧会话日志
会话日志为多帧 zstd(本例父会话 50 帧),
zstdDecompressSync只解第一帧。需按 magic
0xFD2FB528逐帧推进:涉及会话 ID
059dd7d9-62cd-4214-933d-7bfbc11406d74876e8fa-3321-451c-b8be-919693eca8a87f5165d0-95ba-4c79-a095-8fd0afde9d30ac723254-b871-498f-a21f-41c71e825d6019a58976-6309-41f6-a80c-b8e3c9d250aa5de216f2-8df7-4411-b115-94400447695922d1ff8f-c27f-4ab9-8587-865f8760ce3aAll reactions