Replies: 4 comments 4 replies
|
Confirmed against current The proposed direction is good, with one refinement: exclude only the provider-specific bare no-body heuristic, while retaining all other pi-ai overflow signals for every provider. In other words:
Also define the non-Cerebras Regression cases should cover:
That last assembled-loop assertion matters because preventing the misleading recovery cycle is the actual user-visible fix, not only changing one error code. |
|
Thank you for the thorough review against master — much appreciated, especially the 413 fallback catch, which I had missed. Adopted your refinement:
Updated the unit test to cover both bare 400 and bare 413 for a non-Cerebras provider. The fork branch is updated: https://github.com/jc77411/deepseek-harness/tree/fix/pi-ai-bare-400-overflow-misclassification The remaining regression cases you listed (Cerebras bare 400/413 → overflow, descriptive errors → overflow, usage-based silent/length overflow, rate-limit exclusions, and the assembled-loop assertion that no compaction is launched for non-overflow cases) are the right coverage set. I cannot run the assembled-loop test locally here (the monorepo install fails on network access), so that last assertion is the one I would most welcome help validating. |
|
Done — added the assembled-loop sibling case exactly as you described. It uses an in-band Assertions:
Verified locally against the full workspace install: the loop test (7 cases) and the Branch updated: https://github.com/jc77411/deepseek-harness/tree/fix/pi-ai-bare-400-overflow-misclassification |
|
Thank you — appreciate the thorough review throughout. Stable revisionThe tip of
Full SHA of HEAD: The two prior commits: Focused test commands# wire mapping: bare 400 / 413 → INVALID_REQUEST, overflow text untouched
pnpm exec vitest run packages/llm/llm-pi-ai/tests/adapter.spec.ts
# assembled loop: INVALID_REQUEST never enters compaction recovery
pnpm exec vitest run packages/compaction/compaction-basic/tests/compaction-loop-repro.spec.tsBoth pass locally (47 adapter + 7 loop). The branch lives at |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Bug report: bare "400/413 (no body)" is mislabeled as context overflow for non-Cerebras providers
English
Summary
When a non-Cerebras OpenAI-compatible provider (e.g. SiliconFlow) rejects a request with a bare
400 status code (no body),dsh-llm-pi-ai'smapStopReasonlabels itCONTEXT_WINDOW_EXCEEDED. The harness then runs its overflow recovery — a compaction whose summarization request re-sends the same failing provider and fails the same way — collapsing the turn into a retry loop that surfaces a fabricated "context overflow" for a failure that was never about context length.Root cause
mapStopReasondelegates overflow detection to pi-ai'sisContextOverflow, which matches a bare400/413 (no body)against the pattern:That pattern exists only for Cerebras, whose gateway reports overflow as an empty HTTP 400/413. But the harness applies the verdict provider-agnostically, so SiliconFlow's empty 400 (which can mean quota, auth, request guard, or a transient failure) is promoted to
CONTEXT_WINDOW_EXCEEDED.Proposed fix
Guard the pi-ai overflow promotion on the model's provider:
A bare 400/413 (no body) from a non-Cerebras provider then falls through to
classifyPiAiError, which routes the 400 toINVALID_REQUEST. Cerebras keeps the overflow promotion, and descriptive overflow text (matched byisContextWindowExceededErrorvia the separateharnessOverflowarm) is unaffected.Ready-to-review implementation
A complete fix (source change + unit test + Agent Note + bilingual README update) is on my fork branch:
https://github.com/jc77411/deepseek-harness/tree/fix/pi-ai-bare-400-overflow-misclassificationThe single behavioral change is in
packages/llm/llm-pi-ai/src/stream.ts.中文
摘要
当非 Cerebras 的 OpenAI 兼容提供方(如 SiliconFlow)以裸
400 status code (no body)拒绝请求时,dsh-llm-pi-ai的mapStopReason会把它标记为CONTEXT_WINDOW_EXCEEDED。Harness 随后运行溢出恢复——compaction 的摘要请求向同一个失败的提供方重发并以同样方式失败——使回合陷入重试循环,把一个与上下文长度无关的失败呈现为编造的"上下文溢出"。根因
mapStopReason把溢出检测委托给 pi-ai 的isContextOverflow,后者用正则匹配裸400/413 (no body):这条正则仅为 Cerebras 存在——它的网关用空的 HTTP 400/413 表示溢出。但 harness 无差别地应用该判定,于是 SiliconFlow 的空 400(可能意味着配额、鉴权、请求守卫或瞬时故障)被升级为
CONTEXT_WINDOW_EXCEEDED。建议的修复
按模型的 provider 守卫 pi-ai 的溢出升级:
非 Cerebras 提供方的裸 400/413(无 body)随后落入
classifyPiAiError,把 400 归类为INVALID_REQUEST。Cerebras 保留溢出升级,描述性溢出文本(由独立的harnessOverflow分支经isContextWindowExceededError匹配)不受影响。可直接评审的实现
完整的修复(源码改动 + 单元测试 + Agent Note + 双语文档更新)在我的 fork 分支:
https://github.com/jc77411/deepseek-harness/tree/fix/pi-ai-bare-400-overflow-misclassification唯一的行为改动在
packages/llm/llm-pi-ai/src/stream.ts。All reactions