Replies: 3 comments
Desktop confirmation: backgrounded tab + silent WebSocket death matches your resync-race analysis (3 failures + 1 control)Environment: Windows 11 25H2, Node.js 24.14.0, dsh 0.1.0-rc.7 (global npm install), desktop Chrome (not mobile/remote). Failures (long background ~1.5h each): the question/approval card was dispatched while the browser tab was backgrounded; on return the UI showed "waiting for answer / Deep diving..." with an empty composer takeover area. The card only appeared after a manual page refresh (reconnect + mux replay of the still-pending question/requested frame). Session log forensics (decoded .zstd):
All three tool calls/results carry Control (short background ~minutes): a 4th card rendered normally on tab return without refresh — so the long background/silent-disconnect window is required, matching the resync race you identified (replay arrives before Verdict: the bug family is confirmed on the desktop client as well; I support your fix and #3020's server-heartbeat approach. Refresh/reconnect remains the only user-side recovery until upstream adopts a fix. |
|
I verified the rc.7 ordering and combined it with the companion liveness failure in #3020. They are independent layers behind the same missing-card symptom: the Host keeps a stable-rpcId question/approval in its pending registry and replays it on mux open; a silent half-open WebSocket can prevent that new generation entirely, while a real reconnect can deliver the replay before |
|
Independent confirmation on a physical iPhone (iOS 26, Chrome, via a reverse proxy), reached separately before I found this thread — I landed on the same root cause and the same fix shape, which I hope is useful corroboration.
Why it reads as a phone-only bug. A desktop tab holds its socket for hours, so Deterministic reproduction. for (const s of sockets) if (s.readyState === WebSocket.OPEN) s.close()Frame trace from a real run (times relative to that stream's open): The host does its part correctly; the state was lost on the client afterwards. One addition to your fix, from testing a version of it — still relevant to anyone backporting it to 0.1.1.x. Clearing on An exact per-generation marker already existed on the wire: the host pushes A related detail worth stating explicitly, since it's the reason "just delete the two lines" was not equivalent: I ran both variants as a local client-side patch against 0.1.1-rc.2 for a couple of days on a phone; questions survived lock/unlock cycles reliably. Upvoted — thanks for the thorough trace and the prepared fix. |
Uh oh!
There was an error while loading. Please reload this page.
Summary
ask_user_question(and approval prompts) can intermittently fail to render their card in the web UI while the tool call stays blocked. The observed symptom is "stuck — no question card, no tool result"; pressing Stop then surfacesuser abort. The host has the question pending and is correctly waiting for a human answer, but the browser never renders the card to answer.I traced this to a reconnect race in the client runtime and prepared a fix with a deterministic regression test. Since external PRs aren't accepted (per CONTRIBUTING.md), I'm filing it here for the team's attention. Full diff on a fork branch: master...kaywow:fix/web-question-reconnect-race
Root cause
On reconnect:
question/requested/approval/requestedframes (same rpcId), which re-mintPendingWaits on the client — beforeonConnectedfires;onConnectedthen drivesSession.resync(), which unconditionally ranthis.pending.clear().That ordering wipes the just-replayed wait, leaving the host blocked on an answer with no client-side wait to render. The two code comments already disagree:
session.tsassumed "clear first, replay re-mints", whileconnection.tsdocuments "reconnect replays flow from stream open, ahead of onConnected".Fix
Move the clear to generation death:
Session.clearPendingInteractions()— drops all pending waits;SessionManager.handleDisconnected()calls it for every instantiated session — this runs before the next generation's mux replay can arrive;Session.resync()no longer clears pending, so it cannot race the replay.Stale waits (resolved while disconnected) are still dropped — now at disconnect instead of reconnect — and still-pending waits are re-minted verbatim by the replay. Fixes both
questionandapprovalwaits (they share the samependingmap). Files touched:packages/client/runtime/src/client/sessions/{session,manager}.ts.Reproduction (deterministic)
In
packages/client/runtime/tests/session.client.spec.ts:question/requestedframe (the reconnect replay);resync();Before the fix this fails with
expected [] to have a length of 1.Verification
packages/client/runtimesuite: 346 tests pass (includes new regression + manager-level tests).pnpm check:ci(ci-primary, Node 22.23.2): 50/52 gates pass. The two failures are unrelated to this change:test:snapshotleaks anode:sqliteExperimentalWarning into stderr on Node 22, andknippassed in isolation (transient during the full run).All reactions