[Bug] Node >=24 default HTTP/2 fetch corrupts large request bodies against api.deepseek.com: TRANSPORT retry storms (TLS bad record mac + ERR_HTTP2_INVALID_SESSION) #4447
Replies: 1 comment
|
#4448 is a sibling root cause worth reading together #4448 (intranet proxy environment: Node Two observations for whoever picks these up:
Also worth noting: both reports only became diagnosable after patching the adapter to log the swallowed |
Uh oh!
There was an error while loading. Please reload this page.
[Bug] Node ≥24 default HTTP/2 fetch corrupts large request bodies against api.deepseek.com → TRANSPORT retry storms (TLS bad record mac + ERR_HTTP2_INVALID_SESSION)
TL;DR
On Node ≥ 24 (undici v7 enables HTTP/2 by default in
fetch), large request bodies (~75 KB+) sent tohttps://api.deepseek.com(Tencent EdgeOne) intermittently arrive corrupted: the server rejects the TLS stream with abad record macalert, undici destroys the H2 session, and every subsequent retry fails instantly withERR_HTTP2_INVALID_SESSIONbecause the destroyed session stays in the pool. Users seeDeepSeek API request to https://api.deepseek.com failed(TRANSPORT) bursts lasting minutes, then spontaneous recovery. Forcing HTTP/1.1 (Agent({ allowH2: false })) eliminates it completely. This is very likely the same root cause as #978.Environment
b150a551, 2026-08-21)deepseek-official, modeldeepseek-v4-flash,reasoningEffort: high,maxTokens: 256000Symptom
Tasks stream normally for a while, then a turn dies with
DeepSeek API request to https://api.deepseek.com failed(codeTRANSPORT). Session-log timeline of one burst (captured with a local diagnostic patch that logs the swallowed error cause chain):Failure windows last minutes (bursts at 09:53 / 09:54 / 10:00 / 10:09 / 10:10 all failed; 10:21 recovered), then heal spontaneously. Note the sub-15 ms failures on retries — too fast for a network round trip.
Client-side network was ruled out: during a failing window, a plain
curlPOST to the same endpoint from the same machine returns HTTP 401 (reachable, auth-only). DNS, system proxy, VPN, IPv4/IPv6 all verified stable.Root cause (error chains captured)
The shipped
LlmErroronly carries the top-level message; the session log never shows the cause. Patchingpackages/llm/llm-deepseek/src/adapter.tsto logerror.causebefore throwing reveals two distinct layers:First failure — TLS record corruption on the H2 upload:
The server detects the corruption (alert received client-side), i.e. the client's TLS records were already damaged on the wire — with TCP checksums passing, this points at the HTTP/2 write path in the client stack (bundled undici v7 H2 + OpenSSL), not the network.
All 5 retries — poisoned H2 session pool:
Each retry reuses the destroyed session and fails instantly without any connection attempt. The multi-minute "outage windows" are just the pool holding the dead session until cleanup, and users re-triggering bursts with manual retries.
Minimal reproduction (no dsh needed)
Node v26.7.0, macOS. Intermittent but frequent; run the H2 case a few times.
Workaround / suggested fix
An explicit HTTP/1.1 dispatcher in the DeepSeek adapter fixes it end-to-end (verified on a long-running session that previously failed every few minutes):
Beyond the workaround, two upstream-worthy angles:
ERR_HTTP2_INVALID_SESSIONand force a fresh connection instead of replaying against a destroyed session.TODO(http): adopt the Cordis HTTP servicein the adapter — shared transport configuration (including explicit protocol/dispatcher policy) would give users a knob for this class of problems.Happy to turn the diagnostic patch (logging swallowed transport causes) into a PR if useful — it made this whole investigation a 5-minute exercise instead of a guessing game.
Related
deepseek-v4-flash), same "LLM round after a tool call fails with TRANSPORT, retries all fail, cause invisible in session log" pattern; headless vs web only changes how often it surfaces.All reactions