Repository navigation
[Bug] "DeepSeek Messages transport failed" is a masked transport timeout — undici's 300 s bodyTimeout silently caps streamIdleTimeoutMs(传输层 300 秒上限被误报为传输失败) #8947
Replies: 2 comments 1 reply
你的机制与两处默认值恰好撞在一起——我把这条数字对齐核到了1. 关键巧合:两个 300 秒我核到本仓自己的流空闲超时默认值就是 300 秒(当前 HEAD = packages/llm/llm-deepseek/src/defaults.ts:4 export const DEFAULT_STREAM_IDLE_TIMEOUT_MS = 300_000
packages/llm/llm-pi-ai/src/config.ts:47 export const DEFAULT_STREAM_IDLE_TIMEOUT_MS = 300_000而 undici 的 2. 但我要如实说一处证据边界我在 这对修法影响很大,建议在报告里点明:
这一段能让你的报告从"现象描述"变成"可执行的修法方向"。 3. 关于"误报"这一层(与你的标题一致)错误文案
4. 你那条"完全离线可复现"很有价值(无需账号/密钥/外网)——这会让维护者能直接跑。请保留。 5. 版本提醒(本轮基线已变,请以此为准)
一条边界我确认的是两处 300 000 ms 默认值的存在与位置、本仓源码里没有 |
|
如果你的电脑同时拥有卡巴斯基和DSH,出现"DeepSeek Messages transport failed"有可能是因为卡巴斯基的https中间人检查导致的,在卡巴斯基中将api.deepseek.com放入受信任的地址中即可解决这个问题 |
Uh oh!
There was an error while loading. Please reload this page.
English
Summary
streamIdleTimeoutMsis documented as "Maximum provider idle time per outstanding stream read"(
dsh-llm-deepseek/README.md:62), and the same README states that stream-idle expiry throwsTIMEOUT(
:118). Raising it above the default300000therefore looks like a supported way to tolerate a longthinking pause.
It is not. The adapter issues a bare
fetch()against the process-global undici dispatcher that DSHinstalls at boot, and neither DSH's proxy-policy dispatcher nor the adapter overrides
bodyTimeout/headersTimeout. undici's default 300 s body timeout therefore fires first, and the resulting errorescapes the adapter's own
TIMEOUTbranch into the catch-all:So a timeout is reported as a transport failure, and
streamIdleTimeoutMs > 300000is an inert knob.Symptom (observable criteria)
DeepSeek Messages transport failed,code = TRANSPORT, whose cause chain isLlmError → TypeError: terminated → BodyTimeoutError(UND_ERR_BODY_TIMEOUT).llm/retry/llm/retry-startedevents withfailure.code = "TRANSPORT"(the default policy retries
TRANSPORTup to 5 times).streamIdleTimeoutMs > 300000.300000(e.g.4000) and the identical silent stream reportsDeepSeek Messages stream idle timeoutwithcode = TIMEOUT. The watchdog path works; the transporttimeout simply wins the race.
Deterministic reproduction (offline, no credentials, ~10 s)
A loopback SSE endpoint answers
200+ onemessage_startframe and then goes silent forever — nosocket reset, no error frame. The real
DeepSeekAdapterfrom the shipped bundle is driven directly, sonothing is stubbed except the provider.
repro-min.mjs(self-contained;extract-asar.mjsis a 30-line asar reader):Actual output (2026-10-06):
With
--raw(undici defaults, nothing shortened) the second run fails after 301 518 ms with the sameTRANSPORT/BodyTimeoutErrorpair, whilestreamIdleTimeoutMsis900000.Scenario matrix from the fuller harness (each row = a different provider behaviour; same driver):
streamIdleTimeoutMsmessage_start, then silenceBodyTimeoutErrorstream idle timeout(MESSAGES_IDLE after 4000ms)errorSERVERerrorframe (rate_limit_error)RATE_LIMITsocket.destroy()after the first frameSocketError UND_ERR_SOCKET) — the legitimate use of this codeABORTEDS4/S5/S7/S8 are controls:
TRANSPORTshould mean a socket-level failure. A timeout is not one.Expected vs Actual
streamIdleTimeoutMsis the maximum provider idle time per outstanding stream read; idle expiry throwsTIMEOUT;TRANSPORTcovers pre-response transport failures only (README.md:62,:118). Setting900000tolerates a 15-minute silent stream.TRANSPORT;TIMEOUTnever sees it;streamIdleTimeoutMsabove300000has no effect.Evidence (code chain)
TIMEOUTbranch —dsh-llm-deepseek/lib/index.js:2129-2132:The watchdog is armed with the configured value (
:2118), default300000(:17), and only recognises itsown
TimeoutReason("MESSAGES_IDLE"). Any other timeout falls through.The request is a bare
fetch()with no dispatcher (:2188); the body is consumed byparseSseat:2221.Transport timeouts are therefore owned by the process-global dispatcher.
That dispatcher is installed by DSH at boot:
dsh/lib/profile-boot-*.js:225→installProxyFromEnvironment()(dsh-http-proxy/lib/index.js:519) →installGlobalProxy()(:415) →createPolicyDispatcher()(:388-397):No
bodyTimeout/headersTimeout/connectTimeoutoverride anywhere, so undici's Client defaults apply:bodyTimeout = headersTimeout = 300e3(undici/lib/dispatcher/client.js:316-317), connect (including proxytunnel)
10e3(undici/lib/core/connect.js:69).TRANSPORTis in the default retryable set —dsh-llm/lib/types/retry-policy.js:12-21(
maxRetries = 5, backoff 500 ms → 10 s).Impact
streamIdleTimeoutMs > 300000does nothing.README and pointing users at proxies, DNS, TLS and VPNs. This is the real cost: the reporter's first
hypothesis was the local proxy chain, and only an offline reproduction showed the label was wrong.
TRANSPORT, up to 5 retries follow, each with its own 300 stransport cap — a single stalled stream can hold a turn for roughly 30 minutes before the error reaches
the user, while the configured 15-minute patience is neither honoured nor visible.
fetch()inherits the 300 s body cap. Whether it surfacesas
TRANSPORTdepends on each provider's own catch classification (worth auditing for the other adapters).Root cause
Confirmed (code + offline reproduction)
so the 300 s cap is decoupled from
streamIdleTimeoutMs.BodyTimeoutErrorreachesgenerate()'s catch asTypeError: terminatedand lands in the catch-all →TRANSPORT.TIMEOUTbranch itself works (S3) — it simply never wins the race.Inferred (no live capture yet)
the SSE stream; DeepSeek sends SSE comment heartbeats (the adapter renews its watchdog via
onComment), so ahealthy stream never triggers it. A scan of every persisted session transcript on the reporting machine found
no
TRANSPORT/llm/retryrecords. The defect does not depend on that: any 5-minute silence necessarilyproduces the misleading error.
Suggested fix
A. Make the transport timeout follow the configuration (preferred):
If traffic must go through DSH's own proxy policy, the parameter has to reach
ProxyAgent/Poolas well(
createPolicyDispatchercurrently forwardspassedunchanged, with no timeout fields).B. Fix the classification (independent of A, both recommended): map transport-layer timeouts to
TIMEOUT:Only then does
README.md:118hold, and users stop debugging their proxies.C. Fallback if A/B are deferred: clamp
streamIdleTimeoutMsat validation time to the effective transportceiling (≤ 300 000) with an explicit warning, so a value that cannot work is not silently accepted.
中文
摘要
streamIdleTimeoutMs的文档定义是"每次未完成流读取的最大 provider 空闲时间"(
dsh-llm-deepseek/README.md:62),同一份 README 还写明"流空闲到期抛TIMEOUT"(:118)。因此把默认的
300000调大,看起来是"容忍长时间思考停顿"的受支持做法。事实并非如此。适配器用的是不带 dispatcher 的裸
fetch(),走的是 DSH 启动时装的进程级 undici dispatcher;DSH 自己的代理策略 dispatcher 与适配器都没有覆写
bodyTimeout/headersTimeout。于是 undici 默认的300 秒 body timeout 先触发,而这个错误没有进入适配器自己的
TIMEOUT分支,而是落进 catch-all:即:超时被报成传输失败,并且
streamIdleTimeoutMs > 300000是一个失效的旋钮。现象(可观测判据)
DeepSeek Messages transport failed,code = TRANSPORT,cause链为LlmError → TypeError: terminated → BodyTimeoutError(UND_ERR_BODY_TIMEOUT)。llm/retry/llm/retry-started事件且failure.code = "TRANSPORT"(默认策略会重试 5 次)。streamIdleTimeoutMs > 300000。300000以下(例如4000),面对完全相同的静默流会报DeepSeek Messages stream idle timeout(code = TIMEOUT)。看门狗那条路径本身是好的,只是被传输层超时抢了先。
确定性复现(离线、无需凭据,约 10 秒)
一个 loopback SSE 端点返回
200+ 一帧message_start后永久静默——不是断链,也不是错误帧。驱动的是打包里的真
DeepSeekAdapter,除了 provider 之外没有任何桩。repro-min.mjs全文见英文部分的代码块(自包含;extract-asar.mjs是一个约 30 行的 asar 读取器)。实际输出(2026-10-06):
加
--raw(完全用 undici 默认值,不做任何缩短)时,第二条在 301 518 ms 后失败,错误对完全相同,而此时
streamIdleTimeoutMs是900000。更完整的对照矩阵(同一驱动,只改 provider 行为):
streamIdleTimeoutMsmessage_start后静默BodyTimeoutErrorstream idle timeout(MESSAGES_IDLE after 4000ms)errorSERVERerror帧(rate_limit_error)RATE_LIMITsocket.destroy()SocketError UND_ERR_SOCKET)——这才是该 code 的正当用法ABORTEDS4/S5/S7/S8 是对照组:
TRANSPORT应当只表示 socket 级失败,而超时不是。期望 vs 实际
streamIdleTimeoutMs是每次未完成流读取的最大 provider 空闲时间;空闲到期抛TIMEOUT;TRANSPORT只覆盖响应前的传输失败(README.md:62、:118)。设成900000就应容忍 15 分钟静默。TRANSPORT;TIMEOUT分支永远看不到它;streamIdleTimeoutMs在300000以上不产生任何效果。证据(代码链)
TIMEOUT分支 ——dsh-llm-deepseek/lib/index.js:2129-2132(代码见英文部分)。:2118),默认300000(:17),且只认自己的TimeoutReason("MESSAGES_IDLE"),任何别的超时都掉出去。fetch()(:2188),响应体在:2221进入parseSse。因此传输层超时由进程级 dispatcher 决定。
dsh/lib/profile-boot-*.js:225→installProxyFromEnvironment()(dsh-http-proxy/lib/index.js:519)→installGlobalProxy()(:415)→createPolicyDispatcher()(:388-397,代码见英文部分)。没有任何超时覆写,于是 undici 的 Client 默认值生效:
bodyTimeout = headersTimeout = 300e3(
undici/lib/dispatcher/client.js:316-317)、建连(含代理隧道)10e3(undici/lib/core/connect.js:69)。TRANSPORT在默认可重试集合内 ——dsh-llm/lib/types/retry-policy.js:12-21(maxRetries = 5,退避 500 ms → 10 s)。影响面
streamIdleTimeoutMs > 300000完全不起作用。这才是真正的成本:报告者的第一反应就是本地代理链,只有离线复现才证明标签是错的。
TRANSPORT,最多还会重试 5 次,每次尝试各自受 300 秒传输上限约束——一次"停滞型"故障最坏可以把一轮卡住约 30 分钟才把错误交给用户;而配置里那 15 分钟容忍度既没被尊重、也不可见。
fetch()的 provider 都继承这个 300 秒 body 上限;是否显示为TRANSPORT取决于各 provider 自己的 catch 分类(其他适配器值得单独核查)。
根因判断
确认(代码 + 离线复现)
streamIdleTimeoutMs解耦。BodyTimeoutError以TypeError: terminated的形式到达generate()的 catch,落进 catch-all →TRANSPORT。TIMEOUT分支本身是好的(S3),只是抢不到。推断(尚无线上直接证据)
心跳(适配器用
onComment给看门狗续期),心跳正常时不会触发。对报告机器的全部已持久化会话记录扫描后未发现
TRANSPORT/llm/retry记录。缺陷不依赖这一点:只要出现 5 分钟静默,这条误导性报错必然出现。建议修复
A. 让传输层超时服从配置(首选):请求处挂显式 dispatcher,用
streamIdleTimeoutMs作为 body 上限(
headersTimeout另设一个合理的小值);若流量必须走 DSH 自己的代理策略,该参数需要一路传到ProxyAgent/Pool(createPolicyDispatcher现在把passed原样透传,没有超时字段)。B. 修正分类(与 A 独立,建议同时做):把传输层超时映射为
TIMEOUT(
UND_ERR_BODY_TIMEOUT/UND_ERR_HEADERS_TIMEOUT/UND_ERR_CONNECT_TIMEOUT,沿cause链判断)。代码见英文部分。只有这样才能让
README.md:118的契约成立,用户也不会再去查代理。C. 兜底(A/B 短期不做时):在校验层把
streamIdleTimeoutMs钳制到有效上限(≤ 300 000)并显式告警,避免"能填但无效"的配置。
All reactions