Repository navigation
[Bug] pi-ai adapter discards the provider's stable error code, so upstream stream interruptions become non-retryable PI_AI_ERROR #8844
Replies: 7 comments 1 reply
你的定位我核到了根上:分类器只收一个字符串,所以 provider 的稳定错误码在结构上就被丢掉了1. 关键的一行签名
function classifyPiAiError(message: string): string {⇒ 它只接收 这不是"正则写得不够全",而是"接口形状决定了只能靠文本猜"。建议在报告里就这样写——它把这条从"请补一个正则"提升为接口层缺陷,可采纳度完全不同。 2. 这已经是本周第二条同因报告同一形态我上一轮刚答复过一条( 建议你在正文里引用这两条,并把诉求写成:把结构化错误信息(provider code / 3. 请补两样
4. 关于"可复现"你这条的复现不依赖真实断线:只要构造一个带该 一条边界我确认的是分类器签名只收 |
|
Adding a confirmed Windows reproduction of this defect, including its effect on active goals. Observed behaviorTen fatal turns: eight disconnects and two explicit provider overloads. None retried; other TIMEOUT/TRANSPORT failures in the same deployment did retry. Minimal offline reproductionPass this JavaScript on stdin to Actual: PI_AI_ERROR twice. Expected: TRANSPORT and transient SERVER/overload, eligible for bounded retries. The real quota guard also returns false for both texts. Verified installed-source boundary
|
|
补充数据(原发帖人):你提的两条判据都已在 0.2.0-rc.2 上实测,另有一条更靠上的新发现会改变修复方案。 0. 结论先行
1. rtsketo 的判据:本机复现(逐字)$ node --input-type=module - <<'EOF'
const { readFileSync } = await import('node:fs');
const R = '<npm-global>/@deepseek-ai/dsh/node_modules/';
const src = readFileSync(R + '@deepseek-ai/dsh-llm-pi-ai/lib/index.js', 'utf8');
const body = src.split('\n').slice(1379, 1391).join('\n');
const classify = new Function('isQuotaExceededError', 'QUOTA_EXCEEDED_CODE',
body + '\nreturn classifyPiAiError;')(() => false, 'QUOTA_EXCEEDED');
console.log(classify('stream error: stream disconnected before completion: stream closed before response.completed'));
console.log(classify('Our servers are currently overloaded. Please try again later.'));
EOF
PI_AI_ERROR
PI_AI_ERROR你的行号切片( 2. 分类矩阵(20 条,全部真实字符串)我在同一条路径上又跑了 18 条(同一函数、同一安装包),把边界钉死:
这条矩阵本身就是"窄回归"的清单:前两条 + rtsketo 的两条应该各自有固定期望值。 3. 你要的「SSE error 帧原文(含字段名)」3.1 代码可证的部分抛错臂(vendored pi-ai, else if (event.type === "error") {
throw new Error(`Error Code ${event.code}: ${event.message}` || "Unknown error");
}⇒ 该臂读的是 { "type": "error", "code": "upstream_stream_read_error", "message": "Upstream response stream was interrupted" }(这是"能证明存在这三个字段",不是"我抓到了 wire 原文"。) 3.2 日志可证的部分(DSH 只留下文本,没有 code)原始记录( {"turn":7,"reason":{"kind":"error","error":{"message":"Error Code upstream_stream_read_error: Upstream response stream was interrupted","code":"PI_AI_ERROR"}}}注意 3.3 一条推断(已标注)持久化的文本里没有出现 output.errorMessage = formatProviderError(normalizeProviderError(error), `${model.provider === "openai" ? "OpenAI" : model.provider} API error`);
3.4 诚实边界本机没有代理/HTTP 层抓包日志,DSH 也不持久化原始 SSE 帧,所以我没法给出 byte 级原文。如果你的 relay 侧有 access log,那个才是权威原文。 4. 新发现:抛错点在 pi-ai,不在 DSH adapter(影响修复方案)
case "error": {
const text = message.errorMessage ?? "pi-ai stream error";
return { kind: "error", failure: { message: text, code: classifyPiAiError(text) } };
}⇒ 原帖 P1 的"给
5. 你的
|
| 项 | 值 |
|---|---|
| 本机 DSH | 0.2.0-rc.2 |
@deepseek-ai/dsh-llm-pi-ai@0.2.1-alpha.1 vs 本机 |
md5 相同 1858e6f63125e8982528e7a0cdb70ebc → classifyPiAiError 逐字相同 |
| npm dist-tags | latest = 0.2.0-rc.2、next = 0.2.0-rc.2、alpha = 0.2.1-alpha.1 |
⇒ 请把"请确认新版是否仍复现"写成"alpha 通道",并可直接引用上面的 md5 —— alpha 上仍未修。
8. 给 rtsketo 的 narrow regressions(可直接抄)
classifyPiAiError('stream error: stream disconnected before completion: stream closed before response.completed')
expected TRANSPORT (currently PI_AI_ERROR)
classifyPiAiError('Our servers are currently overloaded. Please try again later.')
expected SERVER (currently PI_AI_ERROR)
classifyPiAiError('Error Code upstream_stream_read_error: Upstream response stream was interrupted')
expected TRANSPORT (currently PI_AI_ERROR)
classifyPiAiError('Error Code upstream_http2_stream_error: Upstream HTTP/2 stream failed')
expected TRANSPORT (currently PI_AI_ERROR)
负面回归(不要被顺手改坏):'invalid_request_error: bad payload' 必须仍然是 INVALID_REQUEST;'401 Unauthorized' 必须仍是 AUTH。
9. 关于「不要 blanket-retry PI_AI_ERROR」——同意
PI_AI_ERROR 是 catch-all(lib/index.js:1390 的 return "PI_AI_ERROR"),它同时承载"上游流中断"和"真正未知的 pi 适配器错误"。把它整体加入可重试集合,等于把未知错误也重试。正确做法是先窄化分类(第 8 节),再让 TRANSPORT 走既有重试路径。原帖"重试发生在 open-step 边界、失败的部分输出不会进入模型"这一条我未独立验证,请在正文里标注为原帖主张。
English summary for @rtsketo
- Your offline reproduction is confirmed on DSH
0.2.0-rc.2(exact output in §1): both strings returnPI_AI_ERROR. - New, load-bearing finding: the throwing site is not in DSH's adapter but in the vendored third-party library —
@earendil-works/pi-ai/dist/api/openai-responses-shared.js:640-641. DSH only receivesmessage.errorMessage(a string) at@deepseek-ai/dsh-llm-pi-ai/lib/index.js:1445-1452. So preserving structured provider classification requires a change in pi-ai (or a different read path in the adapter); addingproviderCodetoLlmFailurealone cannot work. - Frame fields: the throwing arm reads
event.type/event.code/event.message; the persisted journal contains only{message, code:"PI_AI_ERROR"}and no provider code field. Raw wire bytes are not persisted by DSH — we cannot supply a captured frame, and we will not invent one. - Your
undefinedtop-level HTTP status claim is corroborated:openai/core/streaming.mjs:45-47throwsnew APIError(undefined, data.error, undefined, response.headers), andformatProviderError(error-body.js:111-117) drops its prefix whenstatus/bodyare undefined — matching the prefix-less text we observe. - Goal impact confirmed:
@deepseek-ai/dsh-goal-round-driver/lib/index.js:200-203disarms onagent/error. Default retryable set is@deepseek-ai/dsh-llm/lib/index.js:251-257(EMPTY_RESPONSE/RATE_LIMIT/SERVER/TIMEOUT/TRANSPORT). - Agreed: no blanket-retrying of
PI_AI_ERROR. Narrow the classification first (§8 gives the exact regression strings), then letTRANSPORTuse the existing retry path. - Version:
dsh-llm-pi-aiis byte-identical between0.2.0-rc.2and0.2.1-alpha.1(md5 1858e6f6…) — not fixed in the alpha channel.
|
追加(同一位发帖人):三条补充,其中第 2 条直接解释你提到的 #8722。 1. 帧的字段集——可以再精确一格除我上一条给的「 export interface ResponseErrorEvent {
code: string | null; // :1980
message: string; // :1984
param: string | null; // :1988
sequence_number: number; // :1992
type: 'error'; // :1996
}即帧是 另外更正一种常见猜测: 2. 两个决定性的 near-miss(其中一个就是 #8722 的机制)
⇒ 你提到的 #8722(Codex 的 (第一、二行的字符串来源: 3. 你的判据可复现性的一处提醒 + 结论落点
|
Supplementary data (original reporter) — English versionEnglish consolidation of my two earlier comments. All commands were run against the installed 1. Your offline reproduction is confirmed on
|
| Input | Returned | Note |
|---|---|---|
Error Code upstream_stream_read_error: Upstream response stream was interrupted |
PI_AI_ERROR | should be TRANSPORT |
Error Code upstream_http2_stream_error: Upstream HTTP/2 stream failed |
PI_AI_ERROR | should be TRANSPORT |
Error Code upstream_stream_truncated: Upstream response stream ended before completion |
TRANSPORT | only by textual coincidence |
stream error: stream disconnected before completion: stream closed before response.completed |
PI_AI_ERROR | your string #1 |
Our servers are currently overloaded. Please try again later. |
PI_AI_ERROR | your string #2 |
overloaded |
PI_AI_ERROR | see §6 — pi-ai's own retry list contains it |
WebSocket stream closed before response.completed |
PI_AI_ERROR | see §6 — this is the #8722 mechanism |
503 Service Unavailable |
SERVER | control: a real status code routes correctly |
500 Internal Server Error |
SERVER | |
429 Too Many Requests / Error Code 429: too many requests |
RATE_LIMIT | |
401 Unauthorized / OpenAI API error (401): Incorrect API key provided |
AUTH | |
insufficient quota |
QUOTA | |
invalid_request_error: bad payload |
INVALID_REQUEST | negative control — must stay |
socket hang up / terminated / fetch failed / read ECONNRESET |
TRANSPORT | |
HTTP2 request did not get a response / other side closed |
TRANSPORT | |
stream ended without message_stop / OpenAI Responses stream ended before a terminal response event |
TRANSPORT | |
pi-ai stream idle timeout after 300000ms |
TIMEOUT |
Sources for the strings (so the matrix is reproducible): @earendil-works/pi-ai/dist/utils/retry.js:22 ("overloaded"), :48 "fetch failed", :54 "socket hang up", :58 "terminated", :66 "stream ended before message_stop"; pi-ai/dist/api/openai-codex-responses.js:1095 ("WebSocket stream closed before response.completed"); pi-ai/dist/api/openai-responses-shared.js:657; dsh-llm-pi-ai/lib/index.js:1910 (idle-timeout template); upstream_* codes and their messages come from the gateway's own Go constants (quoted third-party source, not observed frames).
3. The SSE error frame — what can and cannot be established
Code-backed field set. The throwing arm reads event.type / event.code / event.message — @earendil-works/pi-ai/dist/api/openai-responses-shared.js:640-641:
else if (event.type === "error") {
throw new Error(`Error Code ${event.code}: ${event.message}` || "Unknown error");
}The SDK's own declaration for that event gives the full field list — openai/resources/responses/responses.d.ts:1976-1997:
export interface ResponseErrorEvent {
code: string | null; // :1980
message: string; // :1984
param: string | null; // :1988
sequence_number: number; // :1992
type: 'error'; // :1996
}⇒ the frame is {"type":"error","code":…,"message":…,"param":…,"sequence_number":…}. type/code/message are proven by the captured message text; param/sequence_number are declared-not-observed — no session log stores raw SSE bytes, so I will not present them as captured.
A shape that cannot have produced our message. A frame of the form {"type":"error","error":{…},"sequence_number":…} is intercepted earlier by the SDK at openai/core/streaming.mjs:45-46 and becomes new APIError(undefined, data.error, undefined, response.headers); APIError.makeMessage then uses only error.message (openai/core/error.mjs:17-23). That would surface as the bare string Upstream response stream was interrupted — without the Error Code prefix. Our capture has the prefix ⇒ the observed frame was the flat {type, code, message} form.
What DSH persisted. turn/end in the failing session, verbatim:
{"turn":7,"reason":{"kind":"error","error":{"message":"Error Code upstream_stream_read_error: Upstream response stream was interrupted","code":"PI_AI_ERROR"}}}error.code is DSH's own PI_AI_ERROR; there is no provider-code field anywhere in the record — the provider's stable code survives only as a substring of message. Corresponding assistant/attempt (same turn/step) is in the same journal. A same-session control turn reads {"kind":"error","error":{"message":"DeepSeek Messages transport failed","code":"TRANSPORT"}}.
4. Precision fix: the throwing site is in the vendored pi-ai, not in DSH's adapter
| Layer | Location | Fact |
|---|---|---|
| vendored pi-ai | @earendil-works/pi-ai/dist/api/openai-responses-shared.js:641 |
interpolates event.code into a plain Error message |
| pi-ai Responses path | dist/api/openai-responses.js:154 |
output.errorMessage = formatProviderError(normalizeProviderError(error), …) |
| DSH adapter | @deepseek-ai/dsh-llm-pi-ai/lib/index.js:1445-1452 |
only message.errorMessage (a string) reaches the classifier |
case "error": {
const text = message.errorMessage ?? "pi-ai stream error";
return { kind: "error", failure: { message: text, code: classifyPiAiError(text) } };
}pi-ai declares it string-only: @earendil-works/pi-ai/dist/types.d.ts:367 errorMessage?: string; — there is no structured errorCode field. grep -c 'Error Code' @deepseek-ai/dsh-llm-pi-ai/lib/index.js → 0.
⇒ The original report's P1 ("add providerCode to LlmFailure and carry it at the throw site") cannot work as written. The throw site is in a third-party library and the adapter never sees event.code. Options: (a) upstream pi-ai preserves code on the error object; (b) DSH's adapter takes a path that yields structured fields instead of only errorMessage; (c) the minimal regex narrowing (P0) remains a stopgap — worth labelling as such.
Precision fix on the claim itself: "a structured provider error code never reaches the classifier" is imprecise. The code does reach it — as unstructured text. What never reaches it is the field.
5. SDK paths: your undefined status claim is corroborated
openai/core/streaming.mjs:45-47 (openai 6.40.0):
if (data && data.error) {
throw new APIError(undefined, data.error, undefined, response.headers);
}The status argument is undefined; the nested data.error survives as err.error/code/param/type (openai/core/error.mjs:5-16), while err.status stays undefined. pi-ai's extractStatus (dist/utils/error-body.js:36-46) probes only statusCode → status → $metadata.httpStatusCode → $response.statusCode, never a nested error object, so it returns undefined; formatProviderError (:111-118) then returns norm.message and drops the provider prefix. Inference from that: the prefix-less text we observe implies no recognizable status/body was recovered on that error object.
For completeness on the other streaming API: the Chat-Completions path flattens the nested status the same way, so neither API recovers it into the classification.
6. Two decisive near-misses (one of them is #8722)
overloaded→PI_AI_ERROR, even though pi-ai's own SDK retry list contains"overloaded"(pi-ai/dist/utils/retry.js:22). Upstream considers it retryable; DSH's classifier cannot recognize it.WebSocket stream closed before response.completed→PI_AI_ERROR, for a very specific reason:dsh-llm-pi-ai/lib/index.js:1389matches\bsocket\b, which cannot match insideWebSocket(no word boundary betweenbandS), and its explicit alternative is the exact stringWebSocket closed unexpectedly. So the socket family is covered only by one exact phrase.
⇒ The #8722 case (Codex WebSocket closed 1012 falling through to the catch-all) and this report share one root cause. Suggested merge: add /websocket/i and /overloaded/i at :1389 in addition to the upstream_stream_* patterns — not only the latter.
7. Reproducibility caveat (so others can re-run it)
Quoting the slice as lines 1379-1391 does not run: :1379 is the closing } of the preceding function (a bare new Function over it throws SyntaxError), and classifyPiAiError depends on module-scope bindings from @deepseek-ai/dsh-llm — QUOTA_EXCEEDED_CODE (dsh-llm/lib/index.js:138) and isQuotaExceededError (:181-183) — so without injecting them every input throws ReferenceError at :1382. Slice from 1380 and inject both helpers; then the two PI_AI_ERROR results appear. Your conclusion stands; only the harness needed the fix.
8. Retry decision and goal impact
- Exact drop point:
@deepseek-ai/dsh-llm-retry/lib/index.js:160} else if (!policy.retryableCodes.includes(failure.code)) return next();
- Retryable set:
@deepseek-ai/dsh-llm/lib/index.js:251-257={EMPTY_RESPONSE, RATE_LIMIT, SERVER, TIMEOUT, TRANSPORT}—PI_AI_ERRORabsent (andQUOTAabsent too). - Same-session A/B in real logs: the
TRANSPORTturn retried 5× withpolicyKey = ["normal",5,["EMPTY_RESPONSE","RATE_LIMIT","SERVER","TIMEOUT","TRANSPORT"],500,10000,0.1], while thisPI_AI_ERRORturn has zerollm/retry/llm/retry-startedframes. So the retry executor works; the classification is what fails. - Goal impact confirmed:
agent/erroris emitted atdsh-agent-loop/lib/index.js:880→dsh-goal-round-driver/lib/index.js:201-202disarm(stateFor(agent))→dsh-goal/lib/index.js:624setActivation(agent.session, "disarmed"). API_AI_ERRORturn therefore stops automatic goal continuation, which matches the "goal still looks active but no work continues" observation.
9. Version check
| Package | 0.2.0-rc.2 vs 0.2.1-alpha.1 |
|---|---|
@deepseek-ai/dsh-llm-pi-ai/lib/index.js |
SHA-256 identical (d46dbc30aeb3414190000e6f8aeaaf994246b5e5c651de07639f4a7e92de37ef); the rc.2 tarball is bit-identical to the installed file |
@deepseek-ai/dsh-llm/lib/index.js |
3 hunks, none touching retry policy; DEFAULT_RETRYABLE_CODES block identical |
⇒ unfixed in 0.2.1-alpha.1. (Tag note: 0.2.1-alpha.1 is on the npm alpha tag only; latest/next are 0.2.0-rc.2.)
10. Narrow regressions to add (copy-ready)
classifyPiAiError('stream error: stream disconnected before completion: stream closed before response.completed')
expected TRANSPORT (currently PI_AI_ERROR)
classifyPiAiError('Our servers are currently overloaded. Please try again later.')
expected SERVER (currently PI_AI_ERROR)
classifyPiAiError('Error Code upstream_stream_read_error: Upstream response stream was interrupted')
expected TRANSPORT (currently PI_AI_ERROR)
classifyPiAiError('Error Code upstream_http2_stream_error: Upstream HTTP/2 stream failed')
expected TRANSPORT (currently PI_AI_ERROR)
classifyPiAiError('WebSocket stream closed before response.completed')
expected TRANSPORT (currently PI_AI_ERROR) ← #8722
classifyPiAiError('overloaded')
expected SERVER (currently PI_AI_ERROR)
Negative controls that must not change: 'invalid_request_error: bad payload' → INVALID_REQUEST; '401 Unauthorized' → AUTH.
11. Agreed: do not blanket-retry PI_AI_ERROR
PI_AI_ERROR is the catch-all (:1390), carrying both "upstream stream interrupted" and genuinely unknown pi-adapter failures. Making it retryable wholesale would retry the unknown class too. Narrow the classification first (§10), then let TRANSPORT use the existing path. The original report's claim that a retry happens at the open-step boundary and that failed partial output never reaches the model is the reporter's claim — I have not independently verified it; please label it as such if you cite it.
你这条更正改变了修复方案——我接受,并且据此把方案改成三层1. 接受更正:抛错点不在 DSH 侧你核到抛错在被 vendored 的第三方库 ⇒ 因此我上一轮那句"给 2. ⇒ 修正后的方案(三层,第 3 层今天就能生效)第 1 层·上游(治本):在 pi-ai 的抛错臂里不要只保留 第 2 层·DSH 侧保留原文:你指出 DSH 不保存 SSE wire 原文 ⇒ 出错后无法事后取证。⇒ 建议在失败路径上保留原始错误帧(哪怕只在 第 3 层·文本匹配兜底(今天就能修):你说帧里有
这一层的价值在于:它不依赖 pi-ai 改动、不依赖新的数据流,而你与 3.
|
Reply: I tested your fork patch against the relay codes in the original report — 3 of your 3 messages fixed, 4 related cases still fall throughThanks for the patch and for sharing it as a fork (noted that upstream does not accept external PRs). I applied your two production substitutions to the installed classifier and re-ran my matrix, so here is an independent check rather than a reading of the diff. 1. What your patch fixes — confirmed, no regressionsI took
So: all three messages from your table classify correctly, precedence is preserved, and the permanent/auth/rate-limit controls are untouched. Your test claims line up with what I measure. 2. What still falls through — 4 cases you may want to add
Suggested narrow additions (same shape as yours):
3. Where the provider code is lost (in case it affects the patch's scope)The throwing site is not in DSH's adapter — it is in the vendored third-party library: For the frame itself, the SDK's own type gives the field set — Where the retry decision actually happens: 4. Your severity point is corroborated in the same sessionSame-session A/B from the captured journals: a 5. VersionYour base is the right one to patch: 6. Agreed on scopeKeeping unknown failures on 回复:我用原报告里的 relay code 实测了你的 fork patch —— 你列的 3 条已修好,另有 4 条相邻情形仍落兜底先谢过补丁,也理解你们那边上游不接受外部 PR、所以以 fork 形式给出。我把你的两处生产代码替换应用到了本机已安装的分类器上重跑矩阵,所以下面是独立复测,而不是读 diff 的结论。 1. 你的补丁修好了什么(确认,且无回归)我从已安装的
即:你表里那三条都归对了,优先级未被破坏,认证/配额/限流/非法请求这些既有分支没有副作用。 2. 仍会落兜底的 4 条(建议一并补)
建议补的三条窄模式(与你的写法同形):
3. provider code 丢在哪一层(关系到补丁的边界)抛错点不在 DSH 的 adapter,而在被 vendored 的第三方库: 帧本身,SDK 的类型声明给出字段集: 重试的落点: 4. 你说的严重度,在同会话里有对照同一会话: 5. 版本你的基线选得对: 6. 范围我同意把未知错误留在 |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Summary
A provider SSE stream returned an
errorframe carrying a stable provider code:upstream_stream_read_error/upstream_http2_stream_errorare transport-level stream interruptions(the relay's own read of its upstream SSE body failed). But the pi-ai adapter throws away
event.codeand keeps only the text:
classifyPiAiError(message)then re-derives a DSH code by regex over that text. The regexes do not coverthe
… stream was interruptedwording, so it falls through to the catch-allPI_AI_ERROR. Since thedefault normal-mode retry policy only retries
EMPTY_RESPONSE / RATE_LIMIT / SERVER / TIMEOUT / TRANSPORT, the failure is not retried, and the≈140 KB of tool-call arguments already streamed in that step are silently discarded — the user only
sees one error line.
The core issue: DSH's own documented rule is "consumers route on the code, never on message text",
yet here a stable provider code is downgraded to text and then re-classified by regex.
Symptom (observable criteria)
All four hold for a failing turn:
turn/endcarries the catch-all code:{"kind":"error","error":{"code":"PI_AI_ERROR","message":"Error Code <provider_code>: <provider message>"}}assistant/message— onlyassistant/attempt(holding the partial stream). The whole step is voided.llm/retry/llm/retry-startedevents at all (outside the eligible set → delegates).Measured in one session: retry events per turn =
{turn 3: 10}, turn 7: 0.Control: in the same session, an earlier
TRANSPORTfailure retried 5 times (llm/retry+llm/retry-started× 5, budget exhausted). The retry executor works; only the classification is wrong.
Deterministic reproduction (no relay or network needed)
1. Fake provider
2. Point a pi-ai route at it
(An unresolvable
apiKeyEnvfails withMISSING_CREDENTIAL, not this bug — add the ref first.)3. Send one prompt
llm/retry+llm/retry-started(transport class → should retry)turn/enderror.code = TRANSPORT(or at least expose the provider code)error.code = PI_AI_ERRORError Code upstream_stream_read_error: Upstream response stream was interrupted(code only inside the text)assistant/messagefor the step4. Control
Replace the error frame with a valid
{"type":"response.completed", …}→ the turn completes normally.Root cause
All snippets below are from
app.asar(the@deepseek-ai/dsh-llm-pi-aistream module; its header comment is@module dsh-llm-pi-ai/stream). Verifiable with:① Throw site: stable code downgraded to text
event.codeonly ever becomes part of a string; nothing structured carries it.② Classifier: regex does not cover the "interrupted" wording
For
"Error Code upstream_stream_read_error: Upstream response stream was interrupted"none of the arms match(no 4xx/5xx digits, no timeout wording, no
stream ended before|without, nonetwork|connection|socket|fetch|ECONN*, noterminated|premature close) → catch-allPI_AI_ERROR.③ Decision point
④ Retry point: a code outside the eligible set delegates
From the
@deepseek-ai/dsh-llm-retrydocs:Default normal-mode eligible set, observed in a live
llm/retryevent'spolicyKey:PI_AI_ERROR∉ that set → no retry; the turn ends.⑤ Confirmation from the session journal
The failing step's
assistant/attemptrecord (decoded multi-frame zstd journal):reasoningblocks opened and closed, all emptyblock-start→tool-call(index 3)usageall zeros →finish{kind:"error", failure:{code:"PI_AI_ERROR", …}}So: a very large tool call (≈140 KB of arguments) was cut off after 2 min 20 s, and the whole step —
including everything already streamed — was discarded.
Scope: more than one code is affected
Two of the three stream-error codes used by a common third-party relay family (the code text is traceable to the
open-source
sub2apiproject, which is how I confirmed the code is not DSH's) are misclassified:
upstream_stream_read_errorPI_AI_ERRORTRANSPORTupstream_http2_stream_errorPI_AI_ERRORTRANSPORTupstream_stream_truncatedTRANSPORT(regex happens to matchstream ended before)TRANSPORTThe third one only classifies correctly by textual coincidence.
Impact
(
openai-responsesand friends), instead of retrying like other transport errors.PI_AI_ERRORbecomes a catch-all, so "genuinely unknown" and "upstream stream interrupted" areindistinguishable in journals/telemetry.
Suggested fix
P0 — minimal regex patch
Verified by replicating the original function:
Retrying
TRANSPORThere is safe:dsh-llm-retryrecovers at the agent loop'sagent/request-errorwaterfall (open-step boundary), re-running the failed step in the same turn, and failed partial output
never reaches the model or derived messages.
P1 — let the stable code through (preferred, structural)
Add an optional
providerCode?: stringtoLlmFailure.Make the throw site carry it:
Change the classifier to
classifyPiAiError(text, providerCode?): map the provider code first,regex only as fallback.
Do the same for other adapters (Anthropic / Gemini / OpenAI-compatible) so their stable codes
map into the DSH taxonomy.
P2 — catch-all must stay diagnosable
Keep the provider's original code (and a truncated message) on the failure, and show both the DSH code
and the provider code in the UI error row, so a
PI_AI_ERRORnever means "unknown origin".P3 — visibility of abandoned partial work
When a step fails mid-stream and already-streamed output is dropped, say so ("attempt abandoned,
≈N KB of streamed output discarded") instead of showing only an error code.
Notes / non-bugs
upstream_stream_read_erroris produced by the relay whenits own upstream SSE read fails. DSH cannot prevent it — this report is only about how DSH handles it.
(b) the provider code preserved, (c) visibility when partial output is discarded.
Evidence appendix
Decoding the multi-frame zstd session journal
session.v4.jsonl.zstdis appended frame-by-frame, andzstdDecompressSync/createZstdDecompressonly decode the first frame. Walk the frames by magic
0xFD2FB528:Facts and how each was verified
messageformatError Code <code>: <message>grep -ao 'throw new Error(...)' app.asar(theevent.type === "error"arm of theopenai-responsesstream)classifyPiAiErrortable and catch-allgrep -ac 'classifyPiAiError'hits instrings -a app.asarpolicyKeyof a livellm/retryevent (value above)PI_AI_ERRORis not retryable@deepseek-ai/dsh-llm-retrydocs, "Failures and recovery"grep -ac 'upstream_stream_read_error' app.asar→ 0sub2api(upstream_stream_read_error/upstream_http2_stream_error/upstream_stream_truncated)Minimum repro requirements: a
api: openai-responsespi-ai route + a mid-stream{"type":"error","code":"upstream_stream_read_error", …}frame (a fake server is enough).All reactions