You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
PI_AI_ERROR misclassified as permanent + provider-fabricated completions + compaction into the real death-zone — three reliability fixes that turned "chat permanently broken" into self-healing sessions
#6743
First — thank you for building and open-sourcing DeepSeek Harness. It's a genuinely
thoughtful design (the append-only session journal and the plugin waterfall model in
particular made everything below possible to diagnose and even work around from the user layer, which says a lot about the architecture). I hit a cluster of failures on
a third-party OpenAI-compatible provider and wanted to share what I found and what
fixed it, in case it's useful — with the caveat that this is one account's data, so
please treat the numbers as a reproduction signal, not a general rate.
I know external PRs aren't being accepted right now, so this is filed as a Discussion
per CONTRIBUTING. Where I could, I solved it without patching the core; where I
couldn't, I noted the minimal core change that would help everyone. I'm happy to
reformat/split these or provide redacted session.v3.jsonl.zstd extracts.
Summary
PI_AI_ERROR isn't in DEFAULT_RETRYABLE_CODES, so a transient provider hiccup
looks like a permanent error and the turn dies with zero retries. Observed over a
72h window: 78 failed attempts carrying the message Unknown error (no error details in response) across 22 sessions. The wire shows the provider returning
HTTP 200 + SSE response.failed with error: null under upstream load — classifyPiAiError() maps that to PI_AI_ERROR, which the default policy treats as
terminal, so the user just sees "This turn failed: Unknown error" and (in my case)
couldn't resume that chat at all.
Provider-fabricated completions (finish stop, empty content, usage 0/0/0) are
mapped to EMPTY_RESPONSE, which frames an infrastructure event as a model
behavior. The model never ran; a content-level label sends people down the wrong
debugging path.
Compaction can itself be launched into the provider's real death-zone, creating a
death spiral ("after it compacted, the chat died" — I suspect this is a common
shape of that report).
Evidence
1. Transient failure classified as terminal. Provider z-ai/glm-5.3-free via an
OpenAI-compatible gateway. Wire shape when it degrades:
error: null → "no error details" → PI_AI_ERROR → not in the retryable default set
→ turn/end { reason: { kind: 'error', error: { code: 'PI_AI_ERROR', … } } }. The
same turns succeed minutes later. Deaths clustered in an overnight maintenance band and
cleared after it. I confirmed the intended-vs-actual split by adding PI_AI_ERROR to
that route's retryPolicy.retryableCodes override — which works, proving the only
problem is the default for this whole class of bare-failed-frame OpenAI-compatible
proxies.
2. Fabricated completions. Attempts recorded chunksSeen=['usage','finish'], finishReason='stop', empty content, input_tokens=0. mapStopReason() yields EMPTY_RESPONSE. That's misleading: 0 tokens in + 0 out + no content means the request
never reached the model.
3. The compaction death-zone (real numbers from my journal, provider configured at a
180k window):
metric (72h)
value
highest successful context surface
212,945 tokens
first high-surface death
170,960 tokens
compaction trigger (0.8 × declared)
144,000 tokens
The one-shot summarization request carries the entire oversized region, so compaction
itself can land past the provider's real wall: I see compaction/start → compaction/end
with no compaction/summary event, the surface never shrinks, and every subsequent
turn fails. (My harness-side mitigation is a hierarchical map-reduce summarization plugin
so no single request exceeds the safe band, plus a cross-route failover plugin — both pure
user-layer, no core edits — but the default behavior is the sharp edge.)
What I'd propose (minimal, taste-left-to-you)
Fail-soft for unknown error classes. For LLM requests (replayable from the
append-only journal), treat an unrecognized code as retryable with bounded maxRetries, keeping a small explicit terminal allowlist (INVALID_REQUEST, AUTH, context-window-exceeded) rather than a retryable allowlist. This is the whole
fix for #1 and makes the gateway resilient to the long tail of proxy error shapes.
Label fabricated completions as transport/server. Detect completed-with-zero-usage
and classify as SERVER (or a dedicated code) with a one-line human message like
"provider returned an empty completion without running the model," so retry + UX treat
it as infrastructure.
Guard compaction against its own request size. Either chunk/hierarchically
summarize so the summary request stays under a fraction of the effective window, or —
after N consecutive start→end cycles with no summary — fall back to structural pruning
(non-LLM) instead of re-sending the same oversized request. And surface compaction/failed
with a reason instead of a silent start→end.
(Stretch, optional) Derive a recommended contextWindow from observed usage
(max successful surface vs first high-surface death) and surface it in diagnostics —
it would have told me my effective window was ~210k / dying at ~171k without me writing
a plugin to notice.
For #1 and #2 I have small, test-passing changes against the packages/llm source
(classifier + retry-policy defaults, plus spec updates) — I can paste the diffs inline or
walk through them here whenever you're taking them again. Thanks for reading, and thanks
again for the project — the extensibility is what let me get this much diagnostic clarity
in the first place.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
First — thank you for building and open-sourcing DeepSeek Harness. It's a genuinely
thoughtful design (the append-only session journal and the plugin waterfall model in
particular made everything below possible to diagnose and even work around from the
user layer, which says a lot about the architecture). I hit a cluster of failures on
a third-party OpenAI-compatible provider and wanted to share what I found and what
fixed it, in case it's useful — with the caveat that this is one account's data, so
please treat the numbers as a reproduction signal, not a general rate.
I know external PRs aren't being accepted right now, so this is filed as a Discussion
per CONTRIBUTING. Where I could, I solved it without patching the core; where I
couldn't, I noted the minimal core change that would help everyone. I'm happy to
reformat/split these or provide redacted
session.v3.jsonl.zstdextracts.Summary
PI_AI_ERRORisn't inDEFAULT_RETRYABLE_CODES, so a transient provider hiccuplooks like a permanent error and the turn dies with zero retries. Observed over a
72h window: 78 failed attempts carrying the message
Unknown error (no error details in response)across 22 sessions. The wire shows the provider returningHTTP 200 + SSE
response.failedwitherror: nullunder upstream load —classifyPiAiError()maps that toPI_AI_ERROR, which the default policy treats asterminal, so the user just sees "This turn failed: Unknown error" and (in my case)
couldn't resume that chat at all.
Provider-fabricated completions (finish
stop, empty content, usage 0/0/0) aremapped to
EMPTY_RESPONSE, which frames an infrastructure event as a modelbehavior. The model never ran; a content-level label sends people down the wrong
debugging path.
Compaction can itself be launched into the provider's real death-zone, creating a
death spiral ("after it compacted, the chat died" — I suspect this is a common
shape of that report).
Evidence
1. Transient failure classified as terminal. Provider
z-ai/glm-5.3-freevia anOpenAI-compatible gateway. Wire shape when it degrades:
error: null→ "no error details" →PI_AI_ERROR→ not in the retryable default set→
turn/end { reason: { kind: 'error', error: { code: 'PI_AI_ERROR', … } } }. Thesame turns succeed minutes later. Deaths clustered in an overnight maintenance band and
cleared after it. I confirmed the intended-vs-actual split by adding
PI_AI_ERRORtothat route's
retryPolicy.retryableCodesoverride — which works, proving the onlyproblem is the default for this whole class of bare-
failed-frame OpenAI-compatibleproxies.
2. Fabricated completions. Attempts recorded
chunksSeen=['usage','finish'],finishReason='stop', empty content,input_tokens=0.mapStopReason()yieldsEMPTY_RESPONSE. That's misleading: 0 tokens in + 0 out + no content means the requestnever reached the model.
3. The compaction death-zone (real numbers from my journal, provider configured at a
180k window):
0.8 × declared)compaction/start→compaction/endcompaction/summaryevent, the surface never shrinks, and every subsequentWhat I'd propose (minimal, taste-left-to-you)
append-only journal), treat an unrecognized code as retryable with bounded
maxRetries, keeping a small explicit terminal allowlist (INVALID_REQUEST,AUTH, context-window-exceeded) rather than a retryable allowlist. This is the wholefix for #1 and makes the gateway resilient to the long tail of proxy error shapes.
and classify as
SERVER(or a dedicated code) with a one-line human message like"provider returned an empty completion without running the model," so retry + UX treat
it as infrastructure.
summarize so the summary request stays under a fraction of the effective window, or —
after N consecutive
start→endcycles with no summary — fall back to structural pruning(non-LLM) instead of re-sending the same oversized request. And surface
compaction/failedwith a reason instead of a silent
start→end.(max successful surface vs first high-surface death) and surface it in diagnostics —
it would have told me my effective window was ~210k / dying at ~171k without me writing
a plugin to notice.
For #1 and #2 I have small, test-passing changes against the
packages/llmsource(classifier + retry-policy defaults, plus spec updates) — I can paste the diffs inline or
walk through them here whenever you're taking them again. Thanks for reading, and thanks
again for the project — the extensibility is what let me get this much diagnostic clarity
in the first place.
All reactions