dsh web freezes (event loop blocked) after sending any message — Windows + Node 22 #4755
Replies: 2 comments
|
(以下为同一问题的中文版完整说明,方便中文用户阅读 / Chinese version of the same report) dsh web 在发送任意消息后卡死(事件循环阻塞)——Windows + Node 22摘要在 环境
复现步骤
预期行为模型应正常(流式)回复,就像观察到的一次成功运行: 实际行为
已收集的诊断信息
疑似区域卡死在第一个 LLM turn 完成之后(或 agent 循环开始第一次工具调度步骤时)。候选原因:
临时绕过方案杀掉卡死的 补充说明
感谢关注——工具本身很好,但这个 bug 目前在 Windows 上让它的聊天功能完全不可用。 |
|
The strongest next artifact would be three full process dumps taken while the HTTP probes are timing out, before the watchdog restarts the Host. The current observations prove a live-process availability failure, but they do not yet distinguish a stable main-thread wait from slow native progress. I checked rc.2 commit
That does not clear any native frame, but it means package presence alone should not select the root cause. I would join port 3080 to the exact PID, record fixed-cadence HTTP probes, then capture three private full dumps five seconds apart: $listener = Get-NetTCPConnection -LocalPort 3080 -State Listen -ErrorAction Stop
$dshPid = $listener.OwningProcess
New-Item -ItemType Directory -Force C:\dsh-hang-dumps | Out-Null
procdump64.exe -accepteula -ma -n 3 -s 5 $dshPid C:\dsh-hang-dumpsThe same main-thread native frame across all three samples is much stronger evidence than one stack. If frames move, correlate each sample with the probe and provider timeline. ProcDump Please keep full dumps private: they can contain credentials, prompts, paths, environment values, and model output. A public follow-up can safely include only sanitized module/symbol/offset repetition and timing. I turned the complete capture, WinDbg comparison, direct-vs-TUN A/B, recovery, and reporting sequence into an independent source-backed runbook: https://sandbaseai.github.io/deepseek-harness-handbook/windows-web-event-loop-hang.html |
Uh oh!
There was an error while loading. Please reload this page.
Summary
After sending any message in the
dsh webUI, the wholedsh webprocess becomes unresponsive: the port keeps listening but every HTTP request (including the page itself) times out. The process never recovers; it must be killed and restarted. This reproduces reliably on a fresh process with no third-party plugins.Environment
0.1.1-rc.2(latest on npm, installed globally vianpm i -g @deepseek-ai/dsh)dsh web(also reproduced with--no-open,--port 3081)api.deepseek.comworks fine (fetchto/chat/completionsreturns 200 in ~850 ms from the same machine)Steps to reproduce
dsh web(server starts fine, UI loads, balance/usage widgets render normally).1+1等于几?/hello) in a session.curl http://127.0.0.1:3080times out. Port 3080 is still LISTENING.Expected behavior
The model replies normally (streamed), like the one successful run observed:
你好,请回复OK两个字→ model answeredOKin ~5 s (this single success happened while a Clash TUN proxy was active; the same request hangs when TUN is absent).Actual behavior
dsh webprocess stays alive (memory ~150–180 MB, no crash), but the event loop is blocked:fetchtoapi.deepseek.com/chat/completionsis 200 in <1 s).Diagnostics collected
dsh webuses streaming (stream: true)fetchviadsh-llm-deepseek(fetchImpl = globalThis.fetch) — the transport itself is fine (verified separately).dsh-session-persistence-jsonl, usesnode:zlibzstd APIs).dsh-subprocess-local/dsh-sandbox-localcontainexecSync/spawnSync/Atomics.waitusages;node-addon-landlock-runis also bundled (native module —npm iprinted safe-delete cleanup warnings for itslinux-arm64variants on Windows).Suspected area
Freeze occurs after the first LLM turn completes (or while the agent loop starts its first tool-scheduling step). Candidates:
Atomics.waitindsh-subprocess-local/dsh-sandbox-localif a worker/subprocess never resolves;node-addon-landlock-runnative module behavior on Windows;node:zlib) during the first post-turn save.Workaround
Kill the frozen
dsh webprocess and restart. (I wrote a keep-alive watchdog that probes the port and auto-restarts on two consecutive failures.)Additional notes
deepseek.comtraffic is routed through the proxy group; once the proxy is gone (system direct), every subsequent message freezes the process. This hint may matter for the scheduler/subprocess path.Thanks for looking into this — the tool is great, this bug makes it unusable on Windows at the moment.
All reactions