Replies: 4 comments
|
|
Thanks for the quick look! One clarification that matches your note: on Windows the spin is kernel-mode (the hot thread shows KernelModeTime ≈ 54s vs UserModeTime ≈ 4s) and it happens inside the |
|
UPDATE 4 (2026-08-24): independent reproduction on dsh 0.1.1-rc.2 + isolation test of the ACL sandbox wrapper Same symptom reproduced on a second machine/setup: Live measurements at freeze time (identical to UPDATE 1/2):
Isolation test of the ACL sandbox wrapper itself (same module
No fix on our side yet; we'll evaluate |
|
UPDATE 5 (2026-08-24): workaround CONFIRMED on 0.1.1-rc.2 + browser-use/Bizneo Workaround successfully verified end-to-end in the setup of UPDATE 4 (dsh 0.1.1-rc.2, web profile, Windows): set DSH_PERMISSION_MODE=danger-full-access
dsh web
Still pending a proper upstream fix for the ACL path on 26200 (this is a workaround, not a solution). If helpful, I can describe the exact seam invocation (dsh-sandbox-local runner-in-argv + ACE materialization) that the UPDATE 4 isolation test points to. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Bug:
dsh webfreezes (HTTP dead, process alive) when the agent executes thepwshtool on WindowsVersion:
@deepseek-ai/dsh@0.1.0-rc.7Environment: Windows 10 (10.0.26200.9168) · Node v22.22.3 · PowerShell 7.6.3 · Chrome 2026
Summary
dsh webbecomes completely unresponsive (the page athttp://127.0.0.1:3080/stops answering even for static assets; the port staysLISTENINGand the process stays alive, but every request times out). The server has to be killed to recover. This happens while the agent is executing apwshtool call.Repro (clean environment, no extensions, no user profile)
Start
dsh web.Open the UI in a clean browser profile (reproduced with Playwright
chromiumheadless, no Chrome extensions loaded).Send this message in the chat:
~10 seconds later
http://127.0.0.1:3080/stops responding. The port remainsLISTENING; the Node process is alive but the event loop appears fully blocked (no HTTP response for >40 s; onlytaskkillrecovers it).A second observed trigger: the agent runs
Invoke-WebRequest(e.g. a DuckDuckGo search) inside apwshtool call; after the call is interrupted/retried, the same freeze occurs. Every incident so far (3/3) followed apwshtool execution.Notes
deepseek-officialprovider and a custom OpenAI-compatible route, so it does not depend on the model provider.0.1.0-rc.7;0.1.0-rc.6fails to boot at all on this machine (launcher produces no output and hangs), so downgrading is not a workaround.Possible cause (from reading the shipped bundles)
@deepseek-ai/dsh-subprocess-local(bundled indshrc.7) performs synchronous subprocess operations on the main thread:spawnSync("taskkill", ["/PID", pid, "/T", "/F"], ...)in the Windows process-tree teardown path;execFileSync(...)per poll tick (e.g.DEFAULT_INTERNALS.exec) insidewhile (treeAlive()) await sleepTick()loops.If any of these synchronous calls stalls while a
pwsh/conhostchild is wedged, the Node event loop freezes and the HTTP server stops responding — matching the observed symptom.Impact
Any user on Windows whose agent calls the
pwshtool (very common for Windows workflows) can hit a hard freeze of the whole harness; work in progress is interrupted (sessions are persisted but the loop dies mid-turn).What I'm doing meanwhile
External watchdog that restarts
dshwhen the port stops responding, plus a diagnostics snapshot at freeze time. Happy to attach the snapshot (child processes, CPU, connections) and the Playwright repro script if useful.UPDATE (2026-08-18): root cause refined — CPU busy-loop in the dsh process, not a blocked teardown
Captured with a live sampler while the freeze was happening (port
3080stillLISTENING, HTTP dead for 170+s, then force-killed):pwsh, noOpenConsole, noconhostunder the dsh PID) — so my earlierspawnSync(taskkill)hypothesis does NOT apply on Windows for this symptom.taskkill /Fresolves it.Conclusion: the event loop is stuck in an infinite (or pathologically long) JavaScript busy loop in
dshrc.7, triggered by the agent/tool/Llm path (reproducible with a simplepwshtool prompt in a clean browser). Machine-specific angle: detected on Windows 11 build 26200 (Insider/2026) with Windows Terminal 1.24 (Microsoft Store) and PowerShell 7.6.3 — but plainspawnandnode-ptyspawns ofpwshwork fine in isolation on the same machine.If needed I can attach the full sampler dump and a live CPU/proc snapshot.
UPDATE 2 (2026-08-18): ROOT CAUSE IDENTIFIED + WORKAROUND
UserModeTime≈ 4.5s vsKernelModeTime≈ 54s on the hot thread; process at ~100% of one core), while the main thread is stuck in uninterruptible native code (V8Debugger.pausecannot interrupt it; CPU profile shows main thread idle).pwshtool call under the Windows ACL sandbox (@deepseek-ai/dsh-sandbox-windows-acl, a koffi/Win32 FFI restricted-token implementation). Thepwshexecutor itself uses plainchild_process.spawn(not ConPTY).DSH_PERMISSION_MODE=danger-full-access(ACL sandbox disabled, tools still work, pwsh at FullLanguage).CreateWellKnownSid(WinLocalLogonSid)fails with ERROR_INVALID_PARAMETER"), suggesting this sandbox has build-26200-specific issues. Please check the ACL-restricted-token spawn path (koffi FFI calls) on 26200.UPDATE 3 (2026-08-18): environment detail + question for maintainers
WindowsSelfHostregistry keys are present — the machine has been on an Insider channel at some point).CreateWellKnownSid(WinLocalLogonSid)failure workarounds).dsh-sandbox-windows-aclrunner/restricted-token path busy-spin in kernel on Windows 11 build 26200? Is there a known Win32 quirk withCreateRestrictedToken/SetNamedSecurityInfoW/ job objects on that build? Would a dsh-side patch (e.g., fall back to unconfined or a different Win32 API path on 26200) be the proper fix, or do you have a repro/knowledge that a newer or different Windows build resolves it?DSH_PERMISSION_MODE=danger-full-access.All reactions