[bug] 上下文超长被误判成普通 API 错误:自动压缩 124 次失败 102 次,会话永久卡死 #7632
Replies: 1 comment
|
Thanks for this report — the chain you traced is the one I get too, line for line, and your suggested fix #1 is the right minimal one. I've implemented it as a mountable plugin so it works on the prerelease you are already running, without waiting for a release. The plugin
npm i @argszero/cordis-plugin-length-stop-overflowthen add it through a bundle patch (this is the package's own - insert:
- id: length-stop-overflow
name: '@argszero/cordis-plugin-length-stop-overflow'It observes the public The rule that addresses your #1On an Your three samples classify on the first gate, with no wording consulted at all:
That is the point of reading the numbers instead of the text: Two properties are deliberate, because this rule must not fire on someone else's failure:
One correction, in your favourYour ③ is right that the text fallback misses this. But the usage path was added to The gap is exactly one conjunct. I installed // Case 2: Silent overflow (z.ai style) - successful but usage exceeds context
if (contextWindow && message.stopReason === "stop") { /* input + cacheRead > contextWindow */ }
// Case 3: Length-stop overflow (Xiaomi MiMo style)
if (contextWindow && message.stopReason === "length" && message.usage.output === 0) {
/* input + cacheRead >= contextWindow * 0.99 */
}Ark gave you (Your reading of case 3 is otherwise identical to mine, including the Honest boundary
Your items 4, 5, 6 and 7 stay core-side and are untouched by this: compaction failure still logs ContextThe same package already covers two neighbouring shapes, both of which reach the same dead end through different evidence: a Thanks again for the session-file forensics — the |
Uh oh!
There was an error while loading. Please reload this page.
Summary
在一个长中文 + 代码会话里,真实 prompt 已经涨到模型窗口的 1.31–1.47 倍,但 harness 完全没有意识到「上下文超长」这件事:
{"failure":{"message":"Response incomplete: length","code":"PI_AI_ERROR"}};agent/pre-step静默吞掉(只打一行 warning),turn 照常继续;agent/request-error的溢出恢复逻辑因为failure.code !== CONTEXT_WINDOW_EXCEEDED_CODE而一次都没触发;PI_AI_ERROR死掉,会话从此无法继续,只能新开。关键点:不是压缩没触发,而是压缩在结构上不可能成功,并且所有兜底路径都被同一个错误分类问题挡住了。
Reproduction
api: openai-responses的 provider。本例:contextWindow。status=incomplete+incomplete_details.reason="length"。agent/pre-step只留下step compaction failed: ...; continuing the turn,上下文永不缩小。Current behavior
会话
session-05e8fec4-7176-47a2-9d5b-e111058f5d93(workspaceE:\mu_ai)的session.v3.jsonl.zstd实测数据:1. 压缩统计
compaction/end事件共 124 次:成功 22 次,失败 102 次。首次失败
2026-09-23T09:08:34.393Z,最后一次失败2026-09-23T13:27:49.072Z。最后一次成功是 seq 3189 /
2026-09-23T12:07:31.970Z—— 之后 3 小时内压缩 再没成功过一次。2. 成功的摘要全部顶到 8192 输出上限
最后一次成功的几次摘要,
outputTokens分别是7732, 8071, 8096, 8138, 8129, 7356, 7915—— 明显是贴着maxTokens默认值 8192 被截断的。3. 主请求的死法完全一样
12:52:06.622Z{"failure":{"message":"Response incomplete: length","code":"PI_AI_ERROR"}}13:09:49.447Z13:27:52.190Z4. 真实 prompt 早就超窗了(provider 自己数的)
5486 / 166528 / 99913671 / 171904 / 167277 / 185472 / 16(
totalTokens == inputTokens + cacheReadTokens + outputTokens已在多个样本上核对。注意
outputTokens只有 16 / 999,所以这不是输出被截断,而是输入侧早就超窗。)也就是说:harness 手里已经有能证明溢出的精确 usage 数字,只是错误分类那条路径完全没有用它们。
根因链
① 触发点(第三方依赖)
@earendil-works/pi-ai@0.85.1dist/api/openai-responses-shared.js:660-688mapStopReason:Ark 返回的是
incomplete_details.reason = "length",不等于"max_output_tokens",于是被映射成
stopReason: "error"+ 消息Response incomplete: length。② pi-ai 自己的溢出检测也因此失效
dist/utils/overflow.js:130-155(本机 0.85.1 实测)stopReason === "error"且消息命中OVERFLOW_PATTERNS→"Response incomplete: length"一条都不命中;stopReason === "length"且usage.output === 0→ 这里 stopReason 是error,且 output 是 16/999,两道门都过不去;isRecoverableLength(:163-165)也要求stopReason === "length",同样失效。length映射对还不够。 case 3 里output === 0这个条件在 0.85.1 里依然存在(上游 PR #7540「允许非零 output」虽已合并,但并未出现在本机这个构建里),
所以本次的
output=16 / 999即使 stopReason 正确也仍会被判成"非溢出"。③ DSH 侧退化成通用错误
@deepseek-ai/dsh-llm-pi-ai/lib/index.js:1389-1443mapStopReason@deepseek-ai/dsh-llm/lib/index.js:152-167isContextWindowExceededError的所有正则(
STRUCTURED_CONTEXT_OVERFLOW/TOO_LARGE_FOR_CONTEXT/EXCEEDS_MODEL_CONTEXT…)都要求消息里出现明确的 “context” 字样,
"Response incomplete: length"一条都不命中。→ 最终
classifyPiAiError(text)给出code: "PI_AI_ERROR"。④ 兜底路径被这个 code 挡住
dsh-compaction-basic/lib/index.js:820-839agent/request-error:failure.code !== CONTEXT_WINDOW_EXCEEDED_CODE就直接next()→ 溢出恢复永不触发;dsh-compaction-basic/lib/index.js:798-811agent/pre-step:压缩失败只
log('step compaction failed: ...; continuing the turn')然后next()→ turn 带着超窗上下文继续跑;dsh-compaction-basic/lib/index.js:873-920compactIfNeeded:thresholdTokens = floor(contextWindow * thresholdRatio),超过重试次数后抛compaction still above threshold after N compaction attempts—— 但此时已经晚了。⑤ 阈值本身就算错了
dsh-token-meter/lib/index.js:16与pi-ai/dist/utils/estimate.js:1都用CHARS_PER_TOKEN = 4。在中文 + 代码场景下低估约 1.37×(
pi-ai的 clamp 反推估计值 ≈126k,而 provider 数出 172k)。于是 preset 里的
thresholdRatio: 0.75→ 估计 98,304 ≈ 真实 134k,压缩触发点已经落在真实窗口之外,必然先溢出再压缩。
⑥ 一旦 surface 超窗,压缩在数学上不可能成功
摘要请求要重放整个可压缩前缀(surface 减去约 20,971 token 的 retain 尾巴),
输入本身还是超窗 → 又是
incomplete/length→ 死循环。这就是 124 次触发里 102 次失败、且成功的那 22 次全在窗口还没被撑爆之前的原因。
Expected behavior
contextWindow时,应当被识别为上下文溢出(CONTEXT_WINDOW_EXCEEDED),而不是通用PI_AI_ERROR;agent/request-error)应当被触发,而不是被code挡在门外;Suggested fixes
按「最小改动 → 收益」排序:
dsh-llm-pi-ai:加 usage 兜底判定(收益最大、改动最小,且不依赖上游)在
mapStopReason里,除了看错误文本,再用 provider 的 usage 判定:若
inputTokens + cacheReadTokens >= 0.99 * contextWindow,无论文本是什么都判为CONTEXT_WINDOW_EXCEEDED。本例的 172k / 185k / 192k 会被立刻正确识别。这条不依赖 provider 的
incomplete_details.reason字符串,所以上游不改也能生效。dsh-llm:让isContextWindowExceededError认识 pi-ai 自己的包装文本Response incomplete: <reason>是 pi-ai 造出来的字符串,DSH 认识它才算闭环。dsh-compaction-basic:820-839:溢出恢复不要只认CONTEXT_WINDOW_EXCEEDED_CODE改为
code === CONTEXT_WINDOW_EXCEEDED_CODE || usageBasedOverflow。dsh-compaction-basic:798-811:压缩失败要「响」至少
next(new Error(...))或发出一个用户可见事件;continuing the turn是这次事故里最贵的 11 个单词。dsh-compaction-basic:72:maxTokens: config.maxTokens ?? 8192默认值过低建议默认取
min(model.maxTokens, 某上限)而不是硬编码 8192(本例模型maxTokens: 16384)。dsh-token-meter:16:CHARS_PER_TOKEN = 4对 CJK 严重低估建议按 CJK 字符占比做加权估计。
阈值用真实 usage 而不是估计值
既然 usage 里已经有精确的
input + cacheRead,压力判断就应该以它为准(或至少取max(估计, 上次真实 usage))。上游现状(重要:不能指望上游)
我在写这份报告前先查了上游,结论是DSH 必须自己兜住:
earendil-works/piPR 桌面端 #7817「fix(ai): treat incomplete reason 'length' as a length stop, not an error」提的正是本报告 ① 的修复(描述里明确点名 Doubao / Volcengine Ark),
但创建当天就被
github-actions[bot]自动关闭且未合并,理由是:该仓库的策略是「issue/PR 先自动关闭,维护者事后挑着 reopen」,所以上游修复的时间不可控。
output === 0条件仍会漏掉本例(见上)。也就是说这次事故需要上游两处同时改才对 DSH 生效。
因此我的建议是:DSH 侧先加 usage 兜底(下面的第 1 条),不要等上游。
同时我会另开一个上游 issue 说明 ①+②,但那只是长期对齐,不应当作为 DSH 的修复前提。
附带一条:官方 OpenAI 也会这样
上游 issue #7052 里有一个和本例几乎同构的数据(官方
openai-responses,272k 窗口):{"usage":{"input":3,"output":264,"cacheRead":267059,"cacheWrite":637,"reasoning":110}, "stopReason":"length"}input + cacheRead = 267,062≈ 98% 窗口,output 极小 —— 和本例「输入早已贴满窗口、输出只有 16/999」是同一个现象,只是官方返回的 reason 是
max_output_tokens(能被 pi-ai 正确识别),Ark 返回的是
length(被误判成 error)。这反过来说明:本例的
reason: "length"极可能同样是输入侧窗口耗尽,而不是输出被截断,所以用 usage 判定溢出是语义上正确的做法。
Environment
0.9.0,内置@deepseek-ai/dsh-*0.1.5-rc.2(Windows 11)volcengine-ark,api: openai-responses,baseURLhttps://ark.cn-beijing.volces.com/api/plan/v3deepseek-v4-1-flash,contextWindow: 131072,maxTokens: 16384@earendil-works/pi-ai@0.85.1standard-70(compaction-basic的thresholdRatio被手工设为 0.75)$DSH_HOME\sessions\<workspace-key>\session-05e8fec4-...\session.v3.jsonl.zstd(多帧 zstd,需按 magic
28 B5 2F FD切帧后逐帧解压)相关讨论 / 上游记录
earendil-works/piPR 桌面端 #7817 —— 修 ① 的 PR,被机器人自动关闭未合并。earendil-works/piPR dsh-plugin-teamflow v0.2.0 — from one sentence of requirements to an accepted delivery (multi-agent pipeline for DeepSeek Harness) #7540(已合并)—— 引入「length stop + usage 接近窗口 ⇒ 视为溢出」的思路,但其「允许非零 output」的放宽未出现在
pi-ai@0.85.1构建中。earendil-works/piissue HTTP 402 (insufficient provider balance) never reaches the terminal QUOTA failure code #7052 —— 官方 OpenAI 上的同构现象(输入 ≈98% 窗口 + 极小 output)。earendil-works/piissue [修复方案] 源码启动 prepare 崩溃:profile 解析锚点未规范化导致 tsx 跳过 paths(fork 分支 + 回归测试 + 端到端验证) #7048 —— 摘要被 token cap 截断后仍被当作有效 checkpoint 持久化。All reactions