[Bug] Session becomes permanently un-resumable after a turn ends in error with an unrecorded tool call #2034
Replies: 4 comments
|
Confirmed - this is the #1959/#1697 family's durable scar: the original crash ( Two things you can do right now
Happy to prepare that repair extension as a cherry-pick-ready branch (same queue as our #1697/#2002/#1997 fixes) - it is the mechanism-level completion of this whole family. |
|
错误结束 + 未记录 tool call → 会话永久不可恢复——和 #1841/#1473 同族(状态不一致家族),rc 期最该修的稳定性问题之一。 临时只能开新会话;官方应修"失败时清理/回滚 tool call 状态"。会话恢复坑 FAQ 有:https://github.com/Electricitysheep/dsh-handbook/blob/main/docs/faq.md |
|
Update — root cause found (environment-side, plus an upstream hazard). The Upstream hazard: a plugin declaring an older |
|
Practical confirmation of the repair direction. We recovered a fully bricked session by hand-writing exactly what your repair path produces: an One caveat for whoever implements it: our first rewrite compressed the whole log into a single zstd frame and the app refused to boot with |
Uh oh!
There was an error while loading. Please reload this page.
Summary
When the harness fails mid tool-execution (e.g. the tool runtime scheduler becomes unavailable and
ctx.tools[TOOL_RUNTIME_SCHEDULER]isundefined, throwingCannot read properties of undefined (reading 'prepare')), the affected turn is closed withreason: { kind: "error" }without atool/resultfor the already-recordedtool/call. From that point on the session can never be resumed: every new message re-sends a transcript that ends with an assistanttool_callsblock that has no matching tool result, and the provider API rejects it with:This persists across app restarts — the session is effectively bricked, while other sessions are unaffected.
Root cause analysis
The crash-recovery repair in
dsh-session(interruptedTurnClosers,lib/types/repair.js) only synthesizesinterrupted-tool-resultevents for an open tail turn:Because the crashed turn was already closed with a
turn/end(reason:error), the pending call is cleared and never recovered. The log then containsassistant/message(with a tool-call block) →tool/call→ (no result) →user/message…, which is not a provider-valid transcript.Additionally,
dsh-agent-loop's tool-call scheduling deliberately does not fabricate synthetic results on scheduler failure (see the docstring in theexecuteToolCallssection: "An internal scheduler failure … rejects with the first failure without fabricating tool results"), so the missing result is never written at crash time either.Suggested fix
interruptedTurnClosers(or add a parallel pass) to also match unmatched tool calls in the last closed turn when that turn ended withreason.kind === "error", appending the sameinterrupted-tool-resultevents the open-turn path generates; orexecuteToolCalls, when a scheduler failure occurs aftertool/callwas durably recorded, append an errortool/resultfor the affected calls before failing the turn.Repro notes
node_modulesfiles were being replaced while the app was running and the hot-reload swapped the context, leavingctx.toolsundefined at step execution time).interrupted-tool-resultevent (same shape the repair path generates:ToolOutcomeUnknownError/TOOL_OUTCOME_UNKNOWN) into the session log right after the danglingtool/call, resequenced the remaining eventseqs, and re-framed the zstd JSONL artifact (first frame must contain exactly the header line). The session then resumes normally.Environment
dsh/dsh-session/dsh-agent-loop0.1.0-rc.6 (deepseek-harness), web profile, Windows.All reactions