[Bug] Stream idle watchdog aborts are misreported as TRANSPORT:terminated, and provider-level streamIdleTimeoutMs never reaches the effective timeout layer
#5811
wilbur1305
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
[Bug] Stream idle watchdog aborts are misreported as
TRANSPORT:terminated, and provider-levelstreamIdleTimeoutMsnever reaches the effective timeout layer / [Bug] 流式空闲看门狗中断被误报为 TRANSPORT:terminated,且 provider 级 streamIdleTimeoutMs 配置未传递到实际生效的超时层@deepseek-ai/dsh0.1.2-rc.1@earendil-works/pi-ai0.84.4lmstudio(OpenAI-compatible endpointhttp://localhost:8000/v1, local 27B GGUF model, ~13-24 tok/s)Summary / 摘要
EN: With a slow local model generating long chains of thought, an agent step's streaming LLM request gets cut after roughly 5.8-6 minutes and auto-retried, with the UI showing
失败原因:terminated(TRANSPORT). After a second-by-second correlation with the LM Studio server log and source analysis, two problems were found: (1) Misclassification — there is strong evidence that at least one interruption was triggered by DSH's own stream idle watchdog (the interruption interval exactly equals the user-configuredstreamIdleTimeoutMs: 1200000plus prefill slack), yet the failure was classified asTRANSPORT:terminatedinstead of theTIMEOUT("pi-ai stream idle timeout after 1200000ms") that the code is supposed to throw. (2) Config not reaching the effective timer —profileOptions()does not forwardstreamIdleTimeoutMs, so the provider-level override (1200000ms) has no effect on a cluster of shorter deaths (~350s) still firing at default magnitude. Combined, the user cannot distinguish "our own timeout" from "the server cut the connection", making diagnosis extremely expensive.中文: 在慢速本地模型(长思维链生成)上,agent 步骤的 LLM 流式请求会在约 5.8-6 分钟被中断并自动重试,UI 显示
失败原因:terminated(TRANSPORT)。经与 LM Studio 服务端日志逐秒比对和源码分析,发现两个问题:(1) 误分类——有确凿证据表明至少一次中断是 DSH 自身的流式空闲看门狗触发(中断间隔精确等于用户配置的streamIdleTimeoutMs: 1200000+ 预填余量),但失败被归类为TRANSPORT:terminated而非代码中本应抛出的TIMEOUT("pi-ai stream idle timeout after 1200000ms")。(2) 配置未达——profileOptions()不转发streamIdleTimeoutMs,provider 级覆盖(1200000ms)对短簇死亡(~350s)不生效,实际生效的超时层仍在按默认量级掐断请求。两者叠加导致用户无法从terminated区分"自家超时"与"对端掐断",排障成本极高。Problem 1: Watchdog abort misclassified as TRANSPORT:terminated / 问题 1:看门狗触发被误分类为 TRANSPORT:terminated
Evidence (the "1240-second" case, signature exactly matches the user's config value) / 证据("1240 秒"案例,签名精确命中用户配置值)
The user's settings.yaml for the lmstudio provider / 用户 settings.yaml 中 lmstudio provider 配置:
Full timeline of that request (attempt 2 of turn3 step3) / 该请求(turn3 step3 第 2 次尝试)的完整时间线:
Client disconnected. Stopping generation...and cancels the task / LM Studio 记录并 cancel taskInterruption time − request acceptance = 1240s = 1200000ms (the user's config value) + 40s. Tokens streamed continuously at 17-22 t/s for the entire window (engine-side print_timing every 3s, no stalls) — there was no 20-minute network silence. This interval corresponds exactly to the configured value and can only be explained by the DSH-side watchdog firing.
中断时刻 − 请求受理 = 1240s = 1200000ms(用户配置值)+ 40s。生成全程 token 以 17-22 t/s 持续流出(引擎侧 print_timing 每 3 秒一条,无停顿),不存在 20 分钟的网络静默。该间隔与配置值精确对应,只能解释为 DSH 侧看门狗触发。
Raw event from the DSH session log (note: failure classified TRANSPORT, message "terminated") / DSH session log 中的原始事件(注意 failure 被归类为 TRANSPORT、message 为 "terminated"):
{"type":"llm/retry","time":1788704752926,"data":{"turn":3,"step":3,"provider":"lmstudio","mode":"normal","policyKey":"[\"normal\",5,[\"EMPTY_RESPONSE\",\"RATE_LIMIT\",\"SERVER\",\"TIMEOUT\",\"TRANSPORT\"],500,10000,0.1]","retry":2,"maxRetries":5,"delayMs":1079.xx,"failure":{"message":"terminated","code":"TRANSPORT"}}}Expected vs actual / 期望行为 vs 实际
dsh-llm-pi-ai/lib/index.js:1816-1818explicitly reclassifies watchdog fires as TIMEOUT / 该处代码明确写了看门狗触发时应当改判 TIMEOUT:In practice the code falls through to
throw error, and the raw error is classified by the regex at:1299-1300/ 实际走了throw error,原始错误被:1299-1300的正则按消息归类:Suspected mechanism / 推测机制
EN: The
idleWatchdog.next()fromdsh-timeout(timer cleared infinally) aborts the upstream signal on timeout; the pi-ai SDK's fetch rejects with undici'sTypeError: terminated; in the race,timeoutOf(watchdog.signal)fails to recognize the abort (e.g., the error becomes visible before the signal's abort state, or escapes via theyield/teardown path), and is then captured by theterminatedregex. Suggestion: unconditionally check the watchdog signal's abort reason before regex classification.中文:
dsh-timeout的idleWatchdog.next()(timer 在finally中 clear)超时后 abort upstream signal;pi-ai SDK 的 fetch 以 undici 的TypeError: terminated拒绝;竞态下timeoutOf(watchdog.signal)未能识别(例如错误先于 signal abort 状态可见、或从yield/teardown 路径冒出),随后被terminated正则捕获。建议在错误进入正则分类前,无条件优先检查watchdog.signal的 abort reason。Problem 2: streamIdleTimeoutMs not forwarded to the effective timeout layer / 问题 2:streamIdleTimeoutMs 配置未传递到实际生效的超时层
dsh-llm-pi-ai/lib/index.js:1584-1592—profileOptions()forwards apiKey / reasoning / thinkingBudgets / cacheRetention / transport / timeoutMs, but notstreamIdleTimeoutMs/ 该函数转发上述字段,唯独不转发streamIdleTimeoutMs:Empirical proof: the short-death cluster ignores the 1200000ms override / 实证:短簇死亡无视 1200000ms 覆盖
Six other interruptions in the same session, all at ~336-356s after the first streamed token, incompatible with the 1200000ms override (tokens streamed continuously during generation; no minutes-long silence) / 同一会话中另有 6 次中断,全部发生在首 token 后 ~336-356 秒,与 1200000ms 覆盖完全不符(生成期间 token 持续流出、无分钟级静默):
EN: (For the 6 cases other than #4, intervals cluster at 336-356s after first token; see
DEFAULT_STREAM_IDLE_TIMEOUT_MS = 3e5at:839, schema default at:969, second normalization at:1024. Static analysis could not locate which layer owns this ~350s timer — could the team identify which timer still fires at default magnitude under a 1200s override?)中文:(除 #4 外的 6 次,间隔聚集在首 token 后 336-356s;相关常量见
:839、schema 默认:969、第二处归一:1024。静态分析未能定位该 ~350s 计时器的确切所在层,烦请确认哪个计时器在 1200s 覆盖下仍按默认量级生效。)LM Studio logged this at the exact same second every time (reproduced verbatim 7×) / 每次死亡 LM Studio 侧同刻日志(一字不差复现 7 次):
Problem 3 (minor): the
timeoutfetch option passed by the pi-ai fork is a silent no-op on Node / 问题 3(次要):pi-ai fork 传给 fetch 的 timeout 选项在 Node 上是静默 no-op@earendil-works/pi-ai/dist/api/openai-completions.js:210:timeoutis a Bun-only fetch extension (cf. upstream oh-my-pi PR #2428 for the same class of problem); Node/undici silently ignores it. IftimeoutMsis meant to work on Node, useAbortSignal.timeout()or a custom dispatcher.timeout是 Bun 的 fetch 扩展(参见上游 oh-my-pi PR #2428 对同类问题的处理),Node/undici 会静默忽略。若期望timeoutMs在 Node 生效,需改用AbortSignal.timeout()或自定义 dispatcher。Problem 4 (UX): the retry panel cannot distinguish failure origins / 问题 4(UX):重试面板无法区分失败来源
EN: The trajectory details panel only shows
重试延迟:Xms / 失败原因:terminated("retry delay / failure reason"), rendering "DSH's own watchdog abort" and "the server actually cut the connection" identically. Suggested: surface the original message, the pre-classification error source (watchdog/transport/server), and the effective timeout value in failure details.中文: Trajectory 详情面板只显示
重试延迟:Xms / 失败原因:terminated,对"DSH 自家看门狗 abort"与"对端真的掐了连接"渲染相同文案。建议 failure 详情透出:原始 message、分类前的错误来源(watchdog/transport/server)、生效的超时配置值。Differential diagnosis (already excluded) / 排除性证据(已做的鉴别诊断)
The following were ruled out as root causes, listed to help triage / 以下因素已排除为根因,供 triage 参考:
:8000/v1/chat/completions(stream, max_tokens 6000) streamed for 296 seconds, 12002 SSE lines, cleandata: [DONE], exit 0.TRANSPORT:terminatedat 150s; the first retry succeeded and the step completed normally — the retry path handles genuine network cuts fine. Note: DSH's fetch (undici default dispatcher) ignores http_proxy env vars; no proxy is in this path.Connection errorbursts (qwen38-mtp / jiunsong) were caused by the service not listening; classification was correct and unrelated to this issue.Reproduction / 复现步骤
streamIdleTimeoutMs: 1200000/ 在 settings.yaml 为该 provider 显式配置。TRANSPORT:terminatedafter ~350s; the server logs a client disconnect at the same second. Keeping single-step generation under 5 minutes (e.g., maxTokens ≤ 4000) avoids the failure / 观察:~350s 后报TRANSPORT:terminated,服务端同刻记录 client disconnect。单步生成压进 5 分钟内则不再复现。Suggested fixes / 期望修复
TIMEOUT("pi-ai stream idle timeout after Xms") / 错误进入正则分类前,优先且无条件地检查 TimeoutReason,保证看门狗触发必报 TIMEOUT。streamIdleTimeoutMsinprofileOptions()(or provide separate first-event/idle timeout knobs) so provider-level overrides govern every effective timer / 转发streamIdleTimeoutMs(或提供独立的首事件/空闲超时配置),确保 provider 级覆盖对全部生效计时器有效。timeoutoption with a Node-equivalent (AbortSignal.timeout()) / 在 Node 环境下等价替换 Bun 专属的 fetchtimeout选项。dsh-issue-stream-timeout-misclassified-as-terminated.md
All reactions