[Bug] [DSH 0.1.7-alpha.2] 默认开启的会话日志遥测字段把请求体撑到 205.87 MB,并永久锁死会话(可复现,附实测数据) #7699
Replies: 9 comments
补充三:三轮对抗性自审 —— 含一处对我们自己前面内容的更正上面那份本地实现是我们自己写的、不是上游代码。我们对它做了三轮对抗性自审(每一轮都审上一轮的修复),共改掉 6 处会损坏内容、撑大请求或判断错误的地方。全部如实记在这里。 先说更正:我们最初在这份报告里列了 4 个被改文件,其中 第 1 轮 · 缺陷 1:
|
| 步骤 | blocks |
|---|---|
| index=1 命中 B | [A] → push B′ → [A, B′](长度 2) |
| index=3 命中 D | 已存在 ⇒ 不重新播种 → push D′ → [A, B′, D′](长度 3) |
收尾 slice(3) |
[D, E] → 最终 [A, B′, D′, D, E] |
C 丢失,D 重复,顺序错乱。 实测:
BEFORE FIX blocks=5 order=[A,B,D,D,E] CORRUPT
duplicated=[D] missing=[C]
AFTER FIX blocks=5 order=[A,B,C,D,E] OK
修法:一次 map 全量映射,顺序天然保留。
第 2 轮 · 缺陷 2:摘录可能比原文更长,于是这一遍把 body 撑大,而标记还在撒谎
摘录标记本身有 68 字符开销,所以块长度落在 head+tail 与 head+tail+68 之间时,摘录反而比原文大:
block= 5121 chars excerpt=5186 delta=+65 *** 变大了 ***
block= 5150 chars excerpt=5187 delta=+37 *** 变大了 ***
整个 pass:message bytes 5205 -> 5244 delta=+39 *** 这一遍把 body 撑大了 ***
标记声称:"...[30 bytes omitted...]" 实际:块增大了 37 字节
三个后果:① 该遍在放大 body 而不是缩小;② 我们写的注释 「the ladder only ever shrinks content」被自己证伪;③ 标记报错了省略量。
修法:if (Buffer.byteLength(excerpt) + 8 >= Buffer.byteLength(block.text)) return block; —— 留给更紧的下一级。
+8 是补一个单位不一致:这个判断比的是原始 UTF-8 字节,而上限量的是 JSON 序列化后的字符串,标记里那两个换行在 JSON 里各占 2 字节。只加保护不加余量时 len=5188 仍会 +1(实测抓到)。
穷举验证(六个台阶各自的危险带,每带 400 个长度):
rung [4096,1024] band 5121..5520 worst +0 rung [ 512, 128] band 641..1040 worst +0
rung [2048, 512] band 2561..2960 worst +0 rung [ 256, 64] band 321..720 worst +0
rung [1024, 256] band 1281..1680 worst +0 rung [ 96, 32] band 129..528 worst +0
tested 2400 cases inflated cases = 0 VERDICT: PASS
第 3 轮 · 缺陷 3:isBodylessStatus() 的判定过宽 —— 我们在另一个位置重建了这个 bug 本身
第一版(错):
function isBodylessStatus(status, type, rawMessage) {
if (status !== 413 && status !== 400) return false;
return typeof rawMessage !== "string" || !/context/i.test(rawMessage);
}函数名叫 isBodylessStatus,但它并不检查有没有 body —— 它检查「消息里有没有出现 context 这个词」。于是任何 400,只要返回了结构化 JSON 错误体、而消息里没写 "context",就会被判成上下文超限:
- 参数非法、工具调用格式错、模型名写错 → 全部变成
CONTEXT_WINDOW_EXCEEDED - 然后
dsh-compaction-basic的if (failure.code !== CONTEXT_WINDOW_EXCEEDED_CODE) return next();⇒ 真实错误被拦截掉,转去压缩 - 用户看到「上下文爆了」,真实原因是参数写错;同时白白摘要掉历史
这正是本 issue 在批判的那个毛病 —— 我们在修复里又造了一遍,只是方向相反。 上游不会这样:上游把 400/413 一律归 INVALID_REQUEST(永不自救)。
修法:
function isBodylessStatus(status, type, rawMessage) {
if (status === 413) return true;
if (status === 400) return typeof rawMessage !== "string";
return false;
}实测行为表(7/7 通过):
| status | 结构化消息 | 判定 | 为什么 |
|---|---|---|---|
| 413 | 无(本次的 openresty HTML) | 超限 | 网关拒收 body,这就是本 bug 的现场 |
| 413 | "Request Entity Too Large" | 超限 | 413 按定义就是体积问题 |
| 400 | 无 | 超限 | 没有任何判别信息可用 |
| 400 | "invalid tool schema" | 不碰 | ← 修的就是这一行 |
| 400 | "maximum context length exceeded" | 参考 | 交给 isContextWindowExceededError() 的正则 |
| 429 / 500 | — | 不碰 | 与体积无关 |
这条比其余五条更有参考价值:错误分类既不能靠状态码放宽(会把无关错误吞成「上下文超限」),也不能靠关键词收紧(会永不恢复)。判据应该是「有没有判别信息」,而不是「像不像」。
第 3 轮 · 缺陷 4:压缩的保留下限会超过总数 ⇒ 恢复静默不发生
const SUMMARIZE_INPUT_CAP_TOKENS = 131072;
// 旧:Math.max(65536, Math.floor(totalTokens * 0.5), totalTokens - SUMMARIZE_INPUT_CAP_TOKENS)
// 新:
function summarizeRetainTokens(totalTokens) {
const bounded = Math.max(Math.floor(totalTokens * 0.5), totalTokens - SUMMARIZE_INPUT_CAP_TOKENS);
return Math.max(1, Math.min(bounded, totalTokens - 1));
}65536 这个固定下限可以大于 totalTokens。本机 profile 里就配了 contextWindow: 32768 的模型 —— 那种会话溢出时保留量被算成 65536 ≫ 3 万 ⇒ 要摘要的内容为 0 ⇒ 区间选不出来 ⇒ 恢复静默不发生,413 直接抛给用户。
| total | 旧 retain | 旧摘要输入 | 新 retain | 新摘要输入 |
|---|---|---|---|---|
| 20,000 | 65,536 | −45,536 | 10,000 | 10,000 |
| 32,768 | 65,536 | −32,768 | 16,384 | 16,384 |
| 65,536 | 65,536 | 0 | 32,768 | 32,768 |
| 131,072 | 65,536 | 65,536 | 65,536 | 65,536 |
| 344,000 | 212,928 | 131,072 | 212,928 | 131,072 |
| 900,000 | 768,928 | 131,072 | 768,928 | 131,072 |
新公式同时满足三条:永远有东西可摘要、摘要输入 ≤ 131,072、大 surface 下保留尾部足够厚。
第 3 轮 · 缺陷 5:标记在第二级之后低报丢失量,最坏差 76 倍
每一级是对上一级的摘录再摘录,所以第一版的标记算的是「上一级中间段」的大小,而不是原始的:
rung [4096,1024] 声称 "14880 bytes omitted" 真实丢失 14880 ✅
rung [2048, 512] 声称 "2630 bytes omitted" 真实丢失 17440 *** 低报 ***
rung [ 96, 32] 声称 "260 bytes omitted" 真实丢失 19872 *** 低报 76 倍 ***
模型看到「只省略了 260 字节」,实际蒸发了 19,872 字节 —— 它以为自己还有全文。 这属于静默的错误信息,与本 issue 的类别同源。
修法:用一个 WeakMap 记住每个块被摘录之前的字节数(键是块对象,所以不会给发出去的载荷加任何字段;每级产生的新对象会把这个值带下去),标记始终对原始计量。修后实测:
rung [4096,1024] claims 14880 truly 14880 OK rung [ 512, 128] claims 19360 truly 19360 OK
rung [2048, 512] claims 17440 truly 17440 OK rung [ 256, 64] claims 19680 truly 19680 OK
rung [1024, 256] claims 18720 truly 18720 OK rung [ 96, 32] claims 19872 truly 19872 OK
(判断「摘录是否值得保留」仍然和当前文本比 —— 那才保证每一遍都相对上一遍缩小;标记则和原始比。两个比较用不同的基准,这一点我们一开始也写错过。)
仍然未改的已知限制(有意为之,不是遗漏)
| 发现 | 性质 | 为什么留着 |
|---|---|---|
accept() 被无条件调用:字段被丢弃了,但我们仍告诉插件「已送达」,水位照常推进 ⇒ 那批遥测事件不再重发 |
语义缺陷 | 改成「丢弃时不推进水位」会让水位卡住,把遥测从「偶尔丢」变成「永久滞后」。哪种更合理应由上游定,我们把这个取舍摆出来 |
8 MiB 不是硬保证:六阶阶梯跑完仍超限就 break,照发超限的 body |
已知限制 | 有意取舍:宁可发一个超限请求,也不让整个回合失败。⇒ 这道防线只挡「扩展字段 bloat」,挡不住「消息面自身超限」 |
| 摘录按 UTF-16 字符切,可能劈开代理对(emoji)→ 产生孤立代理项 | 低概率边界 | 建议改法:text.slice(0, head).replace(/[\uD800-\uDBFF]$/, ""),尾部同理。我们的场景碰不到,改法应由上游定 |
每个请求多 2 次全量 JSON.stringify(量字节、base、payload 各一次,原本 1 次) |
性能回退 | ~1.2 MB 载荷下可测;可优化成「量一次、复用」,但会动到 serialize() 的契约 |
| 压力压缩路径会抬高配置的保留量(0.16 → 0.5) | 行为变更 | 这是「把摘要输入封顶」的必然代价,已在代码注释里写明;但用户的 retainRatio 配置在此时不生效 |
我们核实过没有问题的地方
| 检查 | 结果 |
|---|---|
| 会话存储会不会被改坏 | 不会 —— 所有降级都作用在 map() 产出的新对象上,磁盘上的会话不变 |
accept() 的调用时机是否被改动 |
没有 —— 原样包装,时机不变。这正是这个闸门安全的根本原因 |
degradeInlinedImages / capInlinedImages 是否有同类错位 |
没有 —— 二者都按索引替换(perMessage.get(blockIndex) ?? block),顺序天然保留 |
| 真实负载 | 补丁在本机连续跑 4 小时以上(本会话全程 + 三个完整 turn),零崩溃 |
补丁现状(三轮自审之后)
| hunk | 行数 | 字节 | |
|---|---|---|---|
| 完整补丁(适配器,相对上游原始文件) | 13 | +463 / −16 | 27,830 |
| 三轮自审修正增量(相对我们上一版) | 3 | +55 / −21 | 5,951 |
| 被撤出的文件 | — | dsh-llm/lib/types/error.js 还原为上游内容(10,511 → 8,717 字节) |
— |
两份补丁都经过:Myers 生成 → 自检逐字节重建 → 另一套独立实现(PowerShell)再重建一次 → 比对 SHA256 与活文件一致。
一句话
上游的最小修法仍然是一行(enabled 默认值改回 false)。我们贴这份三轮自审,是因为一个声称自己没有缺陷的补丁,通常只是没被审过 —— 六条里有两条会真的写坏内容、一条会把无关的 400 报成上下文超限、一条会让恢复静默失效、一条会撑大请求,还有一条干脆是死代码。其中「死代码」那条是对我们前面内容的更正。
|
下面是我们本机那份字节防线的完整改动(相对上游原始文件)。 两点说明,避免误会:
规模:13 个 hunk,+463 / −16 行,27830 字节。 diff --git a/node_modules/@deepseek-ai/dsh-llm-deepseek/lib/index.js b/node_modules/@deepseek-ai/dsh-llm-deepseek/lib/index.js
--- a/node_modules/@deepseek-ai/dsh-llm-deepseek/lib/index.js
+++ b/node_modules/@deepseek-ai/dsh-llm-deepseek/lib/index.js
@@ -4,7 +4,7 @@
import { getOrCreateAnonymousUserId } from "@deepseek-ai/dsh-anonymous-user-id";
import { MAX_TIMER_DELAY_MS, deadline, idleWatchdog, timeoutOf } from "@deepseek-ai/dsh-timeout";
import { createHash } from "node:crypto";
-import { mkdir, readFile } from "node:fs/promises";
+import { appendFile, mkdir, readFile } from "node:fs/promises";
import { dirname, join } from "node:path";
import { withFileLock, writeFileAtomic } from "@deepseek-ai/dsh-atomic-write";
import { resolveDshHome } from "@deepseek-ai/dsh-home-paths";
@@ -13,6 +13,48 @@
import z from "@deepseek-ai/schemastery";
import { isVolatile } from "@deepseek-ai/cosmokit";
import { credentialRef } from "@deepseek-ai/dsh-credentials";
+// LOCAL PATCH (operator box, 2026-09-24): durable sink for the [DSH-*] diagnostics.
+//
+// Why: the interactive `web` host runs as a CHILD of the GUI launcher
+// (verified: node pid 13516, parent DshWebWindow.exe), so its stderr is not
+// captured anywhere. Every [DSH-REQ] / [DSH-BYTEGUARD] / [DSH-413-CAPTURE] line
+// was therefore lost for interactive sessions and only visible in the
+// `--profile headless` runs that we happened to redirect by hand. That made the
+// 206 MB defect invisible until it had already bricked a session.
+//
+// This tees ONLY lines containing "[DSH-" into a size-capped file, delegating
+// every write to the original implementation, so it cannot change logging
+// behaviour or fail a request.
+//
+// NOTE: the first revision of this patch encoded the path as \u9ecf (= "黏")
+// instead of \u7c98 (= "粘"). The resulting path did not exist, every append
+// failed ENOENT, and the rejection was swallowed -> the sink stayed silent and
+// looked healthy. A silent diagnostic sink is worse than none, so the first
+// failure is now reported loudly through the original stderr writer.
+// Path: D:\粘液整合包 (2)4\粘液整合包\_updates\_tmp\dsh-413\requests.log
+let diagSinkBytes = 0;
+let diagSinkFailed = false;
+const DIAG_SINK_MAX_BYTES = 32 * 1024 * 1024;
+const DIAG_SINK_PATH = "D:\\\u7c98\u6db2\u6574\u5408\u5305 (2)4\\\u7c98\u6db2\u6574\u5408\u5305\\_updates\\_tmp\\dsh-413\\requests.log";
+const diagSinkOriginalWrite = process.stderr.write.bind(process.stderr);
+process.stderr.write = (chunk, encoding, callback) => {
+ try {
+ const diagText = typeof chunk === "string" ? chunk : Buffer.from(chunk).toString("utf8");
+ if (diagText.includes("[DSH-") && diagSinkBytes < DIAG_SINK_MAX_BYTES) {
+ diagSinkBytes += diagText.length;
+ void appendFile(DIAG_SINK_PATH, diagText, "utf8").catch((diagSinkError) => {
+ if (diagSinkFailed) return;
+ diagSinkFailed = true;
+ try {
+ diagSinkOriginalWrite("[DSH-SINK-FAILED] " + DIAG_SINK_PATH + " :: " +
+ (diagSinkError === null || diagSinkError === void 0 ? void 0 : diagSinkError.code) + " " +
+ (diagSinkError === null || diagSinkError === void 0 ? void 0 : diagSinkError.message) + "\n");
+ } catch (_diagSinkReport) {}
+ });
+ }
+ } catch (_diagSinkFailure) {}
+ return diagSinkOriginalWrite(chunk, encoding, callback);
+};
//#region lib/types/model-info.js
/** Protocol-independent model capabilities and reasoning choices. */
const OFF_REASONING_EFFORT = ReasoningEffortId("off");
@@ -901,7 +943,7 @@
* @param prepare - contributor registry captured for this adapter.
* @returns HTTP payload and a commit to invoke only after a successful HTTP response.
*/
-async function prepareRequestExtensions(body, options, prepare) {
+async function prepareRequestExtensions(body, options, prepare, connection) {
let extensions;
try {
extensions = await prepare({
@@ -912,11 +954,49 @@
throw new LlmError("DeepSeek request extension preparation failed", "REQUEST_EXTENSION", { cause: error });
}
for (const field of Object.keys(extensions.fields)) if (Object.hasOwn(body, field)) throw new LlmError(`DeepSeek request extension field ${JSON.stringify(field)} collides with the base request`, "REQUEST_EXTENSION");
+ // LOCAL PATCH (operator box, 2026-09-24): the message-level byte ladder cannot
+ // bound a top-level request field, so this second gate measures every
+ // contributed field and drops the ones that would push the dispatched payload
+ // past the same ceiling. Acceptance is untouched: only the payload shrinks.
+ const ceiling = connection?.maxRequestBodyBytes ?? MAX_REQUEST_BODY_BYTES;
+ const base = JSON.stringify(body);
+ const baseBytes = Buffer.byteLength(base);
+ const fieldBytes = Object.create(null);
+ const kept = Object.create(null);
+ const dropped = [];
+ for (const [field, value] of Object.entries(extensions.fields)) {
+ const size = Buffer.byteLength(JSON.stringify(value));
+ fieldBytes[field] = size;
+ if (baseBytes + size > ceiling - 4096) {
+ dropped.push(field);
+ continue;
+ }
+ kept[field] = value;
+ }
+ const payload = dropped.length === 0 ? JSON.stringify({
+ ...body,
+ ...extensions.fields
+ }) : JSON.stringify({
+ ...body,
+ ...kept
+ });
+ if (dropped.length > 0) try {
+ console.error(`${BODY_GUARD_LOG}${JSON.stringify({
+ t: (/* @__PURE__ */ new Date()).toISOString(),
+ model: options.model,
+ purpose: options.purpose ?? null,
+ ceiling,
+ baseBytes,
+ fieldBytes,
+ dropped,
+ payloadBytes: Buffer.byteLength(payload)
+ })}`);
+ } catch (_logFailure) {}
return {
- payload: JSON.stringify({
- ...body,
- ...extensions.fields
- }),
+ payload,
+ baseBytes,
+ fieldBytes,
+ droppedFields: dropped,
async accept() {
try {
await extensions.accept();
@@ -1170,6 +1250,276 @@
countQuantum: connection.imageOffloadCountQuantum
};
}
+/**
+* LOCAL PATCH (operator box, 2026-09-24): exact dispatched size of one inlined
+* image. `serialize` emits `Buffer.from(version.data).toString("base64")`, so the
+* accounted length is the padded base64 length of the prepared request bytes.
+* @param version - prepared request image, absent when the reference was dropped.
+* @returns base64 byte length, or undefined when the version cannot be measured.
+*/
+function inlineImageBase64Bytes(version) {
+ const bytes = version?.data?.byteLength ?? version?.bytes;
+ return Number.isSafeInteger(bytes) && bytes >= 0 ? Math.ceil(bytes / 3) * 4 : void 0;
+}
+/**
+* LOCAL PATCH (operator box, 2026-09-24): request-wide ceiling on inlined images.
+* The existing knob is a ceiling, not an assertion: the newest occurrences are
+* kept (matching upstream's oldest-first removal policy) and every occurrence
+* that no longer fits becomes the same deterministic text placeholder an
+* `offloaded` occurrence gets. A single image larger than the whole budget is
+* degraded too, and this never throws, so the request still goes out as text.
+* @param messages - image-projected history about to be serialized.
+* @param versions - prepared request versions keyed by attachment id.
+* @param connection - resolved budgets of this route.
+* @returns capped history plus the accounting of one inline decision.
+*/
+function capInlinedImages(messages, versions, connection) {
+ const configured = connection.maxTotalInlineImageBytes ?? MAX_TOTAL_INLINE_IMAGE_BYTES_PER_REQUEST;
+ const single = connection.maxInlineRequestImageBytes ?? MAX_TOTAL_INLINE_IMAGE_BYTES_PER_REQUEST;
+ const budget = Math.max(0, Math.min(configured, single));
+ const maxImages = Number.isSafeInteger(connection.maxImagesPerRequest) ? connection.maxImagesPerRequest : Number.MAX_SAFE_INTEGER;
+ const occurrences = [];
+ for (const [messageIndex, message] of messages.entries()) {
+ if (!Array.isArray(message.content)) continue;
+ for (const [blockIndex, block] of message.content.entries()) {
+ if (block?.type !== "image") continue;
+ occurrences.push({
+ messageIndex,
+ blockIndex,
+ attachment: block.attachment,
+ bytes: inlineImageBase64Bytes(versions.get(block.attachment?.attachmentId))
+ });
+ }
+ }
+ if (occurrences.length === 0) return {
+ messages,
+ inlineImages: 0,
+ degradedImages: 0,
+ inlineBytes: 0
+ };
+ let remaining = budget;
+ let inlineImages = 0;
+ const degraded = /* @__PURE__ */ new Set();
+ for (let index = occurrences.length - 1; index >= 0; index -= 1) {
+ const occurrence = occurrences[index];
+ if (occurrence.bytes !== void 0 && occurrence.bytes <= remaining && inlineImages < maxImages) {
+ remaining -= occurrence.bytes;
+ inlineImages += 1;
+ continue;
+ }
+ degraded.add(occurrence);
+ }
+ if (degraded.size === 0) return {
+ messages,
+ inlineImages,
+ degradedImages: 0,
+ inlineBytes: budget - remaining
+ };
+ const replacements = /* @__PURE__ */ new Map();
+ for (const occurrence of degraded) {
+ const perMessage = replacements.get(occurrence.messageIndex) ?? /* @__PURE__ */ new Map();
+ perMessage.set(occurrence.blockIndex, {
+ type: "text",
+ text: textOnlyImageText(occurrence.attachment)
+ });
+ replacements.set(occurrence.messageIndex, perMessage);
+ }
+ return {
+ messages: messages.map((message, messageIndex) => {
+ const perMessage = replacements.get(messageIndex);
+ if (perMessage === void 0) return message;
+ return {
+ ...message,
+ content: message.content.map((block, blockIndex) => perMessage.get(blockIndex) ?? block)
+ };
+ }),
+ inlineImages,
+ degradedImages: degraded.size,
+ inlineBytes: budget - remaining
+ };
+}
+/**
+* LOCAL PATCH (operator box, 2026-09-24): replace every remaining inline image
+* with its text placeholder. Used by the byte budget, which must bound the body
+* even when the image cap alone is not what makes it too large.
+* @param messages - history about to be serialized.
+* @returns degraded history and how many occurrences changed.
+*/
+function degradeInlinedImages(messages) {
+ let changed = 0;
+ const next = messages.map((message) => {
+ if (!Array.isArray(message.content) || !message.content.some((block) => block?.type === "image")) return message;
+ return {
+ ...message,
+ content: message.content.map((block) => {
+ if (block?.type !== "image") return block;
+ changed += 1;
+ return {
+ type: "text",
+ text: textOnlyImageText(block.attachment)
+ };
+ })
+ };
+ });
+ return {
+ messages: next,
+ changed
+ };
+}
+/**
+* LOCAL PATCH (operator box, 2026-09-24): byte length of the pre-degradation text
+* of every block this guard has already excerpted. Keyed by the block object, so
+* nothing extra reaches the dispatched payload, and the value survives the
+* replacement object each pass creates. Without it the marker would report the
+* PREVIOUS excerpt's middle and under-report the real loss -- measured at 76x on
+* the last rung (claimed 260 bytes omitted when 19,872 were gone).
+*/
+const BODY_GUARD_ORIGINAL_BYTES = new WeakMap();
+/**
+* LOCAL PATCH (operator box, 2026-09-24): excerpt one text block, keeping its
+* head and tail and naming the omitted byte count in the middle. The count is
+* measured against the ORIGINAL text, so a block walked down several rungs keeps
+* reporting the true total loss instead of the last hop's.
+* @param text - original block text.
+* @param headChars - leading characters to keep.
+* @param tailChars - trailing characters to keep.
+* @param originalBytes - byte length of this block before any excerpting.
+* @returns the excerpted text.
+*/
+function bodyGuardExcerpt(text, headChars, tailChars, originalBytes) {
+ const head = text.slice(0, headChars);
+ const tail = text.slice(text.length - tailChars);
+ const total = originalBytes ?? Buffer.byteLength(text);
+ const omitted = Math.max(0, total - Buffer.byteLength(head) - Buffer.byteLength(tail));
+ return `${head}${BODY_GUARD_MARKER(omitted)}${tail}`;
+}
+/**
+* LOCAL PATCH (operator box, 2026-09-24): replace every text block longer than
+* this pass's keep budget with a head+tail excerpt. Never mutates its input.
+* @param messages - current history.
+* @param headChars - leading characters to keep per block.
+* @param tailChars - trailing characters to keep per block.
+* @returns degraded history, how many blocks changed, and characters removed.
+*/
+function truncateTextBlocks(messages, headChars, tailChars) {
+ let changed = 0;
+ let removed = 0;
+ const next = messages.map((message) => {
+ if (!Array.isArray(message.content)) return message;
+ let replaced = false;
+ const content = message.content.map((block) => {
+ if (block?.type !== "text" || typeof block.text !== "string" || block.text.length <= headChars + tailChars) return block;
+ const originalBytes = BODY_GUARD_ORIGINAL_BYTES.get(block) ?? Buffer.byteLength(block.text);
+ const excerpt = bodyGuardExcerpt(block.text, headChars, tailChars, originalBytes);
+ // The marker costs ~68 characters, so a block only just over the keep
+ // budget would come back LARGER than it went in: the pass would inflate
+ // the body and the marker would misreport the omitted count. Leave such
+ // blocks for a tighter rung instead; this is what makes the "only ever
+ // shrinks" property below actually true. The comparison is against the
+ // CURRENT text, so every pass shrinks relative to the previous one.
+ //
+ // The margin covers the difference between the two units: this compares
+ // raw UTF-8 bytes, while the ceiling is measured on the JSON string,
+ // where the marker's two newlines each cost two bytes.
+ if (Buffer.byteLength(excerpt) + 8 >= Buffer.byteLength(block.text)) return block;
+ removed += block.text.length - excerpt.length;
+ changed += 1;
+ replaced = true;
+ const next = {
+ ...block,
+ text: excerpt
+ };
+ BODY_GUARD_ORIGINAL_BYTES.set(next, originalBytes);
+ return next;
+ });
+ if (!replaced) return message;
+ return {
+ ...message,
+ content
+ };
+ });
+ return {
+ messages: next,
+ changed,
+ removed
+ };
+}
+/**
+* LOCAL PATCH (operator box, 2026-09-24): serialize one request under a hard byte
+* ceiling. The dispatched body is the exact JSON string fetch receives, measured
+* after every degradation pass, and the ladder degrades in this order:
+* (1) inline images to their text placeholders, (2) every text block longer than
+* the pass budget to a head+tail excerpt, halving the budget per pass. It always
+* returns a body: a request that cannot fit is still dispatched rather than
+* failing the turn, and one stderr line records what the ladder did.
+* @param options - provider-neutral request.
+* @param connection - validated defaults and byte budgets.
+* @param messages - history to serialize.
+* @param versions - request versions for retained images.
+* @param access - execution-world paths for image descriptions.
+* @param onReplayDegrade - diagnostic for discarded native replay metadata.
+* @param fileIds - resolved Files references; omission selects inline image bytes.
+* @param inline - whether this attempt serializes inline image bytes.
+* @returns the serialized body and the accounting of this request.
+*/
+function fitRequestBody(options, connection, messages, versions, access, onReplayDegrade, fileIds, inline) {
+ const ceiling = connection.maxRequestBodyBytes ?? MAX_REQUEST_BODY_BYTES;
+ const stats = {
+ ceiling,
+ before: 0,
+ bytes: 0,
+ images: 0,
+ texts: 0,
+ passes: 0
+ };
+ let current = messages;
+ let body;
+ let imagesDone = !inline;
+ let ladder = 0;
+ for (let pass = 0; pass < 16; pass += 1) {
+ body = serialize(options, connection, current, versions, access, onReplayDegrade, fileIds);
+ const bytes = Buffer.byteLength(JSON.stringify(body));
+ if (pass === 0) stats.before = bytes;
+ stats.bytes = bytes;
+ if (bytes <= ceiling) break;
+ if (!imagesDone) {
+ imagesDone = true;
+ const step = degradeInlinedImages(current);
+ if (step.changed > 0) {
+ current = step.messages;
+ stats.images += step.changed;
+ stats.passes += 1;
+ continue;
+ }
+ }
+ const keep = BODY_GUARD_EXCERPTS[ladder];
+ if (keep === void 0) break;
+ ladder += 1;
+ const step = truncateTextBlocks(current, keep[0], keep[1]);
+ if (step.changed === 0) continue;
+ current = step.messages;
+ stats.texts += step.changed;
+ stats.passes += 1;
+ }
+ if (stats.passes > 0) try {
+ console.error(`${BODY_GUARD_LOG}${JSON.stringify({
+ t: (/* @__PURE__ */ new Date()).toISOString(),
+ model: options.model,
+ purpose: options.purpose ?? null,
+ ceiling: stats.ceiling,
+ beforeBytes: stats.before,
+ afterBytes: stats.bytes,
+ passes: stats.passes,
+ imagesDegraded: stats.images,
+ textBlocksExcerpted: stats.texts,
+ inline
+ })}`);
+ } catch (_logFailure) {}
+ return {
+ body,
+ ...stats
+ };
+}
function* imageRefs(blocks) {
for (const block of blocks) if (block.type === "image") yield block.attachment;
}
@@ -1482,15 +1832,22 @@
}
/**
* LOCAL PATCH (operator box, 2026-09-24) -- see _updates\_patches\dsh\.
-* A 413 (or bodyless 400) whose rendered message text names no context bound is
-* this provider refusing the request BODY as too large. The only body that grows
-* in this deployment is the message surface, so route it to
-* CONTEXT_WINDOW_EXCEEDED and let compaction recovery actually run instead of
-* the turn dying as INVALID_REQUEST.
+* A 413 is this gateway refusing the request BODY as too large, with or without a
+* structured provider message, so it always routes to CONTEXT_WINDOW_EXCEEDED and
+* lets compaction recovery run instead of the turn dying as INVALID_REQUEST.
+*
+* A 400 is only treated that way when the provider said NOTHING structured at all
+* (rawMessage absent) -- which is what "bodyless" in the name is supposed to mean.
+* An earlier revision also accepted any 400 whose message merely failed to contain
+* the word "context", which re-created this very defect in the opposite direction:
+* a malformed tool call or a rejected parameter would be reported to the user as a
+* context overflow, and compaction would summarise away history chasing a size
+* problem that did not exist.
*/
function isBodylessStatus(status, type, rawMessage) {
- if (status !== 413 && status !== 400) return false;
- return typeof rawMessage !== "string" || !/context/i.test(rawMessage);
+ if (status === 413) return true;
+ if (status === 400) return typeof rawMessage !== "string";
+ return false;
}
/** Classify a provider error without trusting arbitrary response fields.
* @param raw - decoded response or in-band error event.
@@ -1925,17 +2282,23 @@
inline = true;
continue;
}
- const extensions = await prepareRequestExtensions(serialize(options, connection, inline ? inlineImages(messages, versions, connection) : messages, versions, this.imageAccess, (reason) => {
+ // LOCAL PATCH (operator box, 2026-09-24): the inline fallback used to
+ // serialize every image with no request-wide byte ceiling. It now caps
+ // the total inlined image payload and then fits the whole body under the
+ // gateway byte ceiling by degrading toward text, never by throwing.
+ const inlinePlan = inline ? capInlinedImages(messages, versions, connection) : void 0;
+ const bodyPlan = fitRequestBody(options, connection, inlinePlan === void 0 ? messages : inlinePlan.messages, versions, this.imageAccess, (reason) => {
this.dependencies.onReplayDegrade?.({
provider: options.provider,
model: options.model,
reason
});
- }, fileIds), {
+ }, fileIds, inline);
+ const extensions = await prepareRequestExtensions(bodyPlan.body, {
signal,
...options.sessionId === void 0 ? {} : { sessionId: String(options.sessionId) },
...options.purpose === void 0 ? {} : { purpose: options.purpose }
- }, this.dependencies.prepareExtensions);
+ }, this.dependencies.prepareExtensions, connection);
signal.throwIfAborted();
const response = await fetch(`${messagesApiRoot(connection.baseURL)}/messages`, {
method: "POST",
@@ -1971,6 +2334,13 @@
maxTokens: options.maxTokens,
bodyBytes: Buffer.byteLength(extensions.payload),
msgCount: messages.length,
+ inline,
+ inlineImages: inlinePlan?.inlineImages ?? 0,
+ degradedImages: inlinePlan?.degradedImages ?? 0,
+ bodyBeforeGuard: bodyPlan.before,
+ textBlocksExcerpted: bodyPlan.texts,
+ guardPasses: bodyPlan.passes,
+ bodyCeiling: bodyPlan.ceiling,
bodyHead: text.slice(0, 1200),
hdrs: Object.fromEntries(response.headers)
});
@@ -1989,6 +2359,34 @@
});
}
await extensions.accept();
+ // LOCAL PATCH (operator box, 2026-09-24): record EVERY dispatched
+ // request, successful ones included, so the real gateway ceiling can be
+ // derived from what actually goes through instead of only from failures.
+ // baseBytes/fieldBytes split the payload by ownership: the gap between
+ // bodyBytes and baseBytes is exactly what the extension fields added.
+ try {
+ console.error("[DSH-REQ] " + JSON.stringify({
+ t: (/* @__PURE__ */ new Date()).toISOString(),
+ status: response.status,
+ purpose: options.purpose ?? null,
+ sessionId: options.sessionId === void 0 ? null : String(options.sessionId),
+ model: options.model,
+ maxTokens: options.maxTokens,
+ bodyBytes: Buffer.byteLength(extensions.payload),
+ baseBytes: extensions.baseBytes,
+ fieldBytes: extensions.fieldBytes,
+ droppedFields: extensions.droppedFields,
+ msgCount: messages.length,
+ inline,
+ inlineImages: inlinePlan?.inlineImages ?? 0,
+ degradedImages: inlinePlan?.degradedImages ?? 0,
+ bodyBeforeGuard: bodyPlan.before,
+ bodyCeiling: bodyPlan.ceiling,
+ guardPasses: bodyPlan.passes,
+ imagesDegradedByGuard: bodyPlan.images,
+ textBlocksExcerpted: bodyPlan.texts
+ }));
+ } catch (_logFailure) {}
if (response.body === null) throw new LlmError("DeepSeek Messages returned no response body", "EMPTY_RESPONSE");
yield* translate(parseSse(response.body, activity), options.model);
return;
@@ -2006,6 +2404,43 @@
const DEFAULT_MAX_TOKENS = 40000; // LOCAL PATCH 2026-09-24: was 256e3; the completion reserve counts against the request input ceiling (450000-40000-65536 = 344464 threshold)
/** Default bound on accumulated base64 image payload after Files API fallback. */
const DEFAULT_MAX_INLINE_REQUEST_IMAGE_BYTES = 20 * 1024 * 1024;
+/**
+* LOCAL PATCH (operator box, 2026-09-24): request-wide ceiling for the base64
+* images of ONE request. Upstream bounds the inline budget by THROWING
+* IMAGE_OFFLOAD_REQUIRED, which leaves the dispatched body unbounded whenever the
+* recovery hook cannot advance (a 240-image session produced a 206 MB body that
+* the CDN gateway answered with an openresty 413). 4 MiB keeps the newest
+* occurrences and turns every further occurrence into its deterministic text
+* placeholder instead.
+*/
+const MAX_TOTAL_INLINE_IMAGE_BYTES_PER_REQUEST = 4 * 1024 * 1024;
+/**
+* LOCAL PATCH (operator box, 2026-09-24): hard ceiling for one dispatched
+* Messages payload, measured on the exact JSON string handed to fetch. The
+* provider gateway rejects an oversized body with a bodyless 413, and this
+* session's size came from the incremental `dsh_session_log` request field (the
+* session-log suffix, 205,780,728 B) plus the message surface (~0.6 MB), not from
+* image bytes, so a token-based bound can never catch it. 8 MiB sits far above a
+* legitimate 450k-token surface and far below the observed failure.
+*/
+const MAX_REQUEST_BODY_BYTES = 8 * 1024 * 1024;
+/**
+* LOCAL PATCH (operator box, 2026-09-24): head/tail keep budget of each
+* degradation pass, applied to every text block longer than the budget. The
+* ladder only ever shrinks content, so the body converges monotonically.
+*/
+const BODY_GUARD_EXCERPTS = [
+ [4096, 1024],
+ [2048, 512],
+ [1024, 256],
+ [512, 128],
+ [256, 64],
+ [96, 32]
+];
+/** LOCAL PATCH (operator box, 2026-09-24): marker replacing an omitted middle. */
+const BODY_GUARD_MARKER = (omitted) => `\n...[${omitted} bytes omitted to fit the gateway body limit]...[truncated]\n`;
+/** LOCAL PATCH (operator box, 2026-09-24): one-line stderr diagnostic prefix. */
+const BODY_GUARD_LOG = "[DSH-BYTEGUARD] ";
/** Deterministic raw-byte removal step. */
const DEFAULT_IMAGE_OFFLOAD_BYTE_QUANTUM = 64 * 1024 * 1024;
/** Deterministic base64-byte removal step after Files API fallback. */
@@ -2074,6 +2509,10 @@
streamIdleTimeoutMs: z.number().min(Number.MIN_VALUE).max(MAX_TIMER_DELAY_MS).default(DEFAULT_STREAM_IDLE_TIMEOUT_MS).volatile(),
maxRequestFilesBytes: z.number().step(1).min(1).default(DEFAULT_MAX_REQUEST_FILES_BYTES).volatile(),
maxInlineRequestImageBytes: z.number().step(1).min(1).default(DEFAULT_MAX_INLINE_REQUEST_IMAGE_BYTES).volatile(),
+ // LOCAL PATCH (operator box, 2026-09-24): request-wide image byte ceiling and
+ // the hard body ceiling the degradation ladder enforces before fetch.
+ maxTotalInlineImageBytes: z.number().step(1).min(1).default(MAX_TOTAL_INLINE_IMAGE_BYTES_PER_REQUEST).volatile(),
+ maxRequestBodyBytes: z.number().step(1).min(1).default(MAX_REQUEST_BODY_BYTES).volatile(),
maxImagesPerRequest: z.number().step(1).min(1).default(600).volatile(),
imageOffloadByteQuantum: z.number().step(1).min(1).default(DEFAULT_IMAGE_OFFLOAD_BYTE_QUANTUM).volatile(),
inlineImageOffloadByteQuantum: z.number().step(1).min(1).default(DEFAULT_INLINE_IMAGE_OFFLOAD_BYTE_QUANTUM).volatile(),
@@ -2147,6 +2586,12 @@
if (!Number.isSafeInteger(maxRequestFilesBytes) || maxRequestFilesBytes <= 0) throw new Error("llm-deepseek: maxRequestFilesBytes must be a positive safe integer");
const maxInlineRequestImageBytes = config.maxInlineRequestImageBytes ?? 20971520;
if (!Number.isSafeInteger(maxInlineRequestImageBytes) || maxInlineRequestImageBytes <= 0) throw new Error("llm-deepseek: maxInlineRequestImageBytes must be a positive safe integer");
+ // LOCAL PATCH (operator box, 2026-09-24): re-judge the two local byte budgets
+ // exactly like every other knob so a programmatic caller cannot skip them.
+ const maxTotalInlineImageBytes = config.maxTotalInlineImageBytes ?? MAX_TOTAL_INLINE_IMAGE_BYTES_PER_REQUEST;
+ if (!Number.isSafeInteger(maxTotalInlineImageBytes) || maxTotalInlineImageBytes <= 0) throw new Error("llm-deepseek: maxTotalInlineImageBytes must be a positive safe integer");
+ const maxRequestBodyBytes = config.maxRequestBodyBytes ?? MAX_REQUEST_BODY_BYTES;
+ if (!Number.isSafeInteger(maxRequestBodyBytes) || maxRequestBodyBytes <= 0) throw new Error("llm-deepseek: maxRequestBodyBytes must be a positive safe integer");
const maxImagesPerRequest = config.maxImagesPerRequest ?? 600;
if (!Number.isSafeInteger(maxImagesPerRequest) || maxImagesPerRequest <= 0) throw new Error("llm-deepseek: maxImagesPerRequest must be a positive safe integer");
const imageOffloadByteQuantum = config.imageOffloadByteQuantum ?? 67108864;
@@ -2182,6 +2627,8 @@
streamIdleTimeoutMs,
maxRequestFilesBytes,
maxInlineRequestImageBytes,
+ maxTotalInlineImageBytes,
+ maxRequestBodyBytes,
maxImagesPerRequest,
imageOffloadByteQuantum,
inlineImageOffloadByteQuantum,
核验链:Myers 生成 → 自检逐字节重建 → 另一套独立实现(PowerShell)再重建一次 → SHA256 与活文件一致。 补丁 SHA256: (另有一份仅含三轮自审修正的增量 diff,见「413自审修正-可粘贴」。) |
|
这是仅含三轮对抗性自审修正的增量 diff —— 即"我们上一版实现 → 修正后实现"之间的全部改动。 改动只有 3 个 hunk,+55 / −21 行,落在 同样注意:这是对打包产物而不是仓库源码的 diff。 diff --git a/node_modules/@deepseek-ai/dsh-llm-deepseek/lib/index.js b/node_modules/@deepseek-ai/dsh-llm-deepseek/lib/index.js
--- a/node_modules/@deepseek-ai/dsh-llm-deepseek/lib/index.js
+++ b/node_modules/@deepseek-ai/dsh-llm-deepseek/lib/index.js
@@ -1367,17 +1367,31 @@
};
}
/**
+* LOCAL PATCH (operator box, 2026-09-24): byte length of the pre-degradation text
+* of every block this guard has already excerpted. Keyed by the block object, so
+* nothing extra reaches the dispatched payload, and the value survives the
+* replacement object each pass creates. Without it the marker would report the
+* PREVIOUS excerpt's middle and under-report the real loss -- measured at 76x on
+* the last rung (claimed 260 bytes omitted when 19,872 were gone).
+*/
+const BODY_GUARD_ORIGINAL_BYTES = new WeakMap();
+/**
* LOCAL PATCH (operator box, 2026-09-24): excerpt one text block, keeping its
-* head and tail and naming the omitted byte count in the middle.
+* head and tail and naming the omitted byte count in the middle. The count is
+* measured against the ORIGINAL text, so a block walked down several rungs keeps
+* reporting the true total loss instead of the last hop's.
* @param text - original block text.
* @param headChars - leading characters to keep.
* @param tailChars - trailing characters to keep.
+* @param originalBytes - byte length of this block before any excerpting.
* @returns the excerpted text.
*/
-function bodyGuardExcerpt(text, headChars, tailChars) {
+function bodyGuardExcerpt(text, headChars, tailChars, originalBytes) {
const head = text.slice(0, headChars);
const tail = text.slice(text.length - tailChars);
- return `${head}${BODY_GUARD_MARKER(Buffer.byteLength(text) - Buffer.byteLength(head) - Buffer.byteLength(tail))}${tail}`;
+ const total = originalBytes ?? Buffer.byteLength(text);
+ const omitted = Math.max(0, total - Buffer.byteLength(head) - Buffer.byteLength(tail));
+ return `${head}${BODY_GUARD_MARKER(omitted)}${tail}`;
}
/**
* LOCAL PATCH (operator box, 2026-09-24): replace every text block longer than
@@ -1392,23 +1406,36 @@
let removed = 0;
const next = messages.map((message) => {
if (!Array.isArray(message.content)) return message;
- let blocks;
- for (const [index, block] of message.content.entries()) {
- if (block?.type !== "text" || typeof block.text !== "string" || block.text.length <= headChars + tailChars) continue;
- blocks ??= message.content.slice(0, index);
- const excerpt = bodyGuardExcerpt(block.text, headChars, tailChars);
+ let replaced = false;
+ const content = message.content.map((block) => {
+ if (block?.type !== "text" || typeof block.text !== "string" || block.text.length <= headChars + tailChars) return block;
+ const originalBytes = BODY_GUARD_ORIGINAL_BYTES.get(block) ?? Buffer.byteLength(block.text);
+ const excerpt = bodyGuardExcerpt(block.text, headChars, tailChars, originalBytes);
+ // The marker costs ~68 characters, so a block only just over the keep
+ // budget would come back LARGER than it went in: the pass would inflate
+ // the body and the marker would misreport the omitted count. Leave such
+ // blocks for a tighter rung instead; this is what makes the "only ever
+ // shrinks" property below actually true. The comparison is against the
+ // CURRENT text, so every pass shrinks relative to the previous one.
+ //
+ // The margin covers the difference between the two units: this compares
+ // raw UTF-8 bytes, while the ceiling is measured on the JSON string,
+ // where the marker's two newlines each cost two bytes.
+ if (Buffer.byteLength(excerpt) + 8 >= Buffer.byteLength(block.text)) return block;
removed += block.text.length - excerpt.length;
changed += 1;
- blocks.push({
+ replaced = true;
+ const next = {
...block,
text: excerpt
- });
- }
- if (blocks === void 0) return message;
- for (const block of message.content.slice(blocks.length)) blocks.push(block);
+ };
+ BODY_GUARD_ORIGINAL_BYTES.set(next, originalBytes);
+ return next;
+ });
+ if (!replaced) return message;
return {
...message,
- content: blocks
+ content
};
});
return {
@@ -1805,15 +1832,22 @@
}
/**
* LOCAL PATCH (operator box, 2026-09-24) -- see _updates\_patches\dsh\.
-* A 413 (or bodyless 400) whose rendered message text names no context bound is
-* this provider refusing the request BODY as too large. The only body that grows
-* in this deployment is the message surface, so route it to
-* CONTEXT_WINDOW_EXCEEDED and let compaction recovery actually run instead of
-* the turn dying as INVALID_REQUEST.
+* A 413 is this gateway refusing the request BODY as too large, with or without a
+* structured provider message, so it always routes to CONTEXT_WINDOW_EXCEEDED and
+* lets compaction recovery run instead of the turn dying as INVALID_REQUEST.
+*
+* A 400 is only treated that way when the provider said NOTHING structured at all
+* (rawMessage absent) -- which is what "bodyless" in the name is supposed to mean.
+* An earlier revision also accepted any 400 whose message merely failed to contain
+* the word "context", which re-created this very defect in the opposite direction:
+* a malformed tool call or a rejected parameter would be reported to the user as a
+* context overflow, and compaction would summarise away history chasing a size
+* problem that did not exist.
*/
function isBodylessStatus(status, type, rawMessage) {
- if (status !== 413 && status !== 400) return false;
- return typeof rawMessage !== "string" || !/context/i.test(rawMessage);
+ if (status === 413) return true;
+ if (status === 400) return typeof rawMessage !== "string";
+ return false;
}
/** Classify a provider error without trusting arbitrary response fields.
* @param raw - decoded response or in-band error event.
核验链:Myers 生成 → 自检逐字节重建 → 另一套独立实现(PowerShell)再重建一次 → SHA256 与活文件一致。 补丁 SHA256: |
|
=== 发到 #7699 的评论:#7699 === 补充四:同族报告清单、我们独占的部分,以及一处更正 写这条,是因为我们后来把 Discussions 翻了一遍,发现同一族问题在 24 小时内至少有 5 份独立报告,而我们这条一条都没有引用。先把这个补上。 同族报告(按首次出现排序)
我们要更正的两处
我们真正独占的三条
顺带:目前现成可用的缓解
|
|
补一条源码定位,把「零水位」这条路径钉在行号上。 本机构建产物里,afterSeq 的 -1 有【两个】来源,而第一个是被注释文档化的常态: lib/index.js:83 @returns greatest accepted sequence, or 完整链条:默认开启(119) → 从未接受过 ⇒ -1(87,注释 83 明说) → offset 0 ⇒ 整份日志(127) 要点:#7658 覆盖的是 :96,我们实测的是 :87。两者都汇到 -1,且 :83 的注释说明 口径:以上为构建产物 lib/index.js 的行号;源码路径行号以 #7658 / argszero 给出的 |
|
Thank you for the measured report — the A. The default flip is one release earlier than the report saysThe report attributes it to -const RECORDED_SESSION_LANE = RECORDED_SESSION_LANE_MARKERS
- enabled: z.boolean().default(!RECORDED_SESSION_LANE),
+ enabled: z.boolean().default(true),The release boundary, read out of the tags rather than the changelog:
One more thing about that same commit: the design note shipped with it concedes the failure mode in advance.
"No prompt tokens are added" is the reason your D2 is not a gap in the token model — it is the token model working as designed on an input it cannot see. And the third sentence is this bug, written down on the day the default changed. B. Mapping your artifact diff back to source pathsYou asked for this in the full-diff comment. On the default branch, the functions you patched live here (all line numbers at
One naming note: the concept you are adding already exists in the tree, but only on the inbound client side — C. Your fix suggestion #5 (
|
|
Three things about the release boundary, since the tree-level reading above and the published artifacts do not say the same thing. The injected field now has a bound, in a published version
const Config = z.object({
enabled: z.boolean().default(true),
maxBytes: z.number().step(1).min(1).default(8 * 1024 * 1024)
});The contribution builds the pending suffix event by event and stops at the first one that would cross the ceiling ( The code also states its own residual case: if the first pending event alone crosses the ceiling, nothing is contributed and it logs Two boundaries on this. First, The 413 classification did not changeIn else if (isContextWindowExceededError(detail)) code = "CONTEXT_WINDOW_EXCEEDED";
else if (status === 400 || status === 413 || type === "invalid_request_error") code = "INVALID_REQUEST";So a bodyless 413 whose detail does not match the context-window predicate still lands on The token axis has no equivalent, and a different endingThe ceiling above is on bytes. On the token side we have seen requests go out above the window rather than be stopped before sending: provider errors reporting a requested size somewhere between roughly 1.2× and about 6× the model's window, each carrying the same wording ( The other ending of the same inflation is not a refused request: The first line is a mark-compact pass reporting a reduce at the several-gigabyte level. Separately, in a reproduction that holds one session in a single process, the heap grew from a few megabytes to several hundred (RSS past a gigabyte), and settled back into the hundreds of megabytes after a GC. Adjacent in time, not in cause: a As an observation rather than a claim about intent: the two axes now end differently. A request that is too large in bytes is refused, and the watermark explains the retry; a request that is too large in tokens is issued, and if the process dies there is no watermark to explain anything. The byte axis has a check before sending in this release; the token axis does not have its counterpart. Authorship note: the release comparisons above were made by me against published tarballs of the versions named; the crash and request readings are from my own runs; drafting assisted by AI; verification and publication by me. I have no engineering background — if any technical claim reads wrong, please call it out; I will re-verify against the toolchain and correct. Reported by the OfferKuai Team — Founder: Zhaofeng (Yaming). Website: https://www.offerkuai.com/ | Contact: contact@offerkuai.com 中文版补三件关于发布边界的事 —— 因为"仓库上的读数"和"已发布产物"说的不是同一件事。 那个注入字段现在有界了,而且是在一个已发布版本里
const Config = z.object({
enabled: z.boolean().default(true),
maxBytes: z.number().step(1).min(1).default(8 * 1024 * 1024)
});贡献逻辑改成逐个事件累积、在第一个会越界的那个事件处停下( 代码自己也写明了剩下的那种情况:如果第一条待发事件自己就跨过上界,就什么都不贡献,并打一条日志 有两处边界要说清。第一, 413 的分类没有变在 else if (isContextWindowExceededError(detail)) code = "CONTEXT_WINDOW_EXCEEDED";
else if (status === 400 || status === 413 || type === "invalid_request_error") code = "INVALID_REQUEST";也就是说,在 token 轴没有等价物,而且结局不一样上面那个上界是字节上的。在 token 这一侧,我们见到的是"请求照样发出去、并没有在发送前被拦住":提供方错误里报出的请求规模落在窗口的约 1.2× 到约 6× 之间,措辞完全一样( 同一种膨胀的另一个结局,不是"请求被拒": 第一行是 mark-compact 那一步报出的归并,量级在数 GB。另一次复算里的 heap 读数:一个进程持着一个会话,heap 从几兆涨到数百兆(RSS 过 G),GC 之后回落到数百兆。时间上相邻、但不是因果:先是 这里只作一个观察,不作关于意图的主张:两个轴的结局现在不一样了。字节上太大的请求会被拒,而接受标记能解释这次重试;token 上太大的请求发了出去,一旦进程死掉,没有任何标记能解释。这个版本里,字节轴有了一道发送前的检查;token 轴还没有它的对应物。 声明:上列发布版本之间的比对由我本人针对已发布产物完成;崩溃与请求读数为我方现场读数;文稿撰写由 AI 辅助;核验与发布由我本人负责。我没有工程背景 —— 若任何技术表述有误,请直接指出,我会对照工具链重新核实并更正。我们是插件作者、不是维护者,本文不陈述任何官方政策。 本报告由 OfferKuai(Offer快)团队提交 —— 创始人:Zhaofeng(Yaming)。官网:https://www.offerkuai.com/ | 联系:contact@offerkuai.com |
|
补充五:更正归属 —— 本问题的首发报告是 #6754,不是我们 我们在动手写之前没有检索 Discussions。刚才补做了,结果是:本报告的几乎每一个核心判断,都已被别人更早、有时更准确地写过了。逐项更正如下。 一、首发是 #6754(@yuanBigRun,2026-09-15),而且它比我们完整
它的实测里已经包含我们在本报告里称为"零水位"的那条发现:
它还引了文档自身的承认(
它给出的修法与绕法与我们一致:给字段加大小上限(超限则跳过且不写水位,或分段+检查点上传);在 唯一的环境差异是报错表面:他们是 macOS / master 对照结论:默认开启、 二、版本归属错了:是 0.1.6,不是 0.1.7-alpha.2由 @argszero 指出,我们复核确认: 0.1.6-alpha.1 / 0.1.6-alpha.2 / 0.1.7-alpha.1 都带着它 —— 受影响面比我们写的更广;反过来 0.1.5-rc.3 是针对这一个缺陷的真实回退目标。我们原文表格里「只有遥测默认值是 alpha 引入的」这一行作废。 @argszero 还引出该提交自带的设计说明
三、"三条共性问题"里有两条不是我们的
我们当时是独立复现,但没有先搜。归属更正给 @PerryLink 与上述各帖。 四、所以我们真正还剩什么(缩小到四条)
五、补一个新测量:rc.2 的修复有效,代价可以量化@yamingmou 指出上游已在 三点观察:
顺带一个被这条测量否决的猜想:我们原本怀疑「单条事件自身超过 8 MiB 会让水位永久停住」(rc.2 代码里写明的残留情形),但实测最大单条事件只有 0.31 MiB(工具结果本身已被截断),所以这个残留情形在正常会话上够不到 —— 我们不写这条。 六、致谢@yuanBigRun(#6754,本问题的首发报告)、@argszero(版本更正、源码路径映射、 本报告里每一个关键判断,都是被别人先说出或纠正之后才成立的。 我们剩下的价值只在第四、五两节。 |
|
Three readings about the version boundary of one mechanism, taken from published artifacts. 1. One of the pitfalls is already handled upstream — and it landed in rc.2, not earlierThe second point in The version boundary is sharp. It is also absent from the master revision used for the source mapping earlier in this thread: at 2. What is missing is the trigger, not the mechanismThe fallback above fires on a failure to serialize. The HTTP error path is separate and drops nothing: on a non-2xx the adapter parses the body, gives 3. Where the rest of that inventory is readThe three adapter-level items grouped under the second guardrail of Boundary: read from published tarballs and registry metadata on 2026-09-25; nothing was executed. The patch text earlier in this thread is its author's and I did not apply it. The source-path mapping is the previous comment's — I re-read artifacts rather than the tree, so where the two disagree, what I report is the artifact. Authorship note: the readings above were taken by me against published npm artifact tarballs; drafting assisted by AI; verification and publication by me. I have no engineering background — if any technical claim reads wrong, please call it out; I will re-verify against the toolchain and correct. Reported by the OfferKuai Team — Founder: Zhaofeng (Yaming). Website: https://www.offerkuai.com/ | Contact: contact@offerkuai.com 中文版三条读法,讲的是一个机制的版本边界,全部取自已发布产物。 1. 其中一条"踩坑",上游其实已经做到了 —— 而且它是 rc.2 才有的
版本边界很干脆: 同样地,上文用来做源码映射的那个 master 版本里也没有它:在 2. 缺的是触发条件,不是机制上面那条兜底的触发条件是序列化失败。HTTP 错误路径是另一条、且什么都不丢:非 2xx 时适配器解析响应体、给 3. 这份清单的其余部分在读在哪里
边界:以上读于 2026-09-25,取自已发布 tarball 与登记处元数据,未执行任何代码。本线程前面那份 patch 文本是它的作者的,我没有应用。源码路径映射是上一条评论的 —— 我读的是产物而不是源码树,所以两者不一致时,我报的是产物。 声明:上列读数由我本人针对已发布 npm 产物 tarball 完成;文稿撰写由 AI 辅助;核验与发布由我本人负责。我没有工程背景 —— 若任何技术表述有误,请直接指出,我会对照工具链重新核实并更正。我们是插件作者、不是维护者,本文不陈述任何官方政策。 本报告由 OfferKuai(Offer快)团队提交 —— 创始人:Zhaofeng(Yaming)。官网:https://www.offerkuai.com/ | 联系:contact@offerkuai.com |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
🚀 【省流版 / TL;DR】
一句话总结:这不是上下文超限,而是默认开启的遥测日志撑爆请求体,并因水位线机制形成永久自锁死循环。
一、 核心真相:0.58 MB 的真实对话 vs 205.87 MB 的遥测垃圾本机实测,真实对话消息面(baseBytes)仅有 0.58 MB。然而,被拒的请求体真实体积高达 206 MB。其中,dsh_session_log 这个遥测字段竟然占到了 205,872,373 字节(205.87 MB),占整个请求体的 99.7%。结论:与用户的上下文长度、图片体积毫无关系,完全是遥测数据在作祟。
二、 为什么会永久死锁?(自锁循环机制)1. 插件 dsh-session-log-deepseek 向每个请求注入 dsh_session_log 字段,内容为“上次被接受水位之后的全部会话事件”。2. 水位线只有在 HTTP 2xx(成功)之后才会推进。3. 因为该字段过大,网关直接返回 413 Request Entity Too Large 拦截。4. 被拒后,水位线不推进,下次请求重发同样的(甚至因为新事件而更大的)载荷。5. 循环成立,会话永久无法推进。
三、 为什么系统“察觉不到、救不回来”?(全版本共有缺陷)· D1(版本归因):dsh-session-log-deepseek 在 0.1.7-alpha.2 中默认值从 false 翻转为 true。两版 README 描述截然相反,这是本 bug 的导火索。· D2(无字节防线):压力模型完全基于 token(默认窗口设到 1e6,触发线 678k)。对图片有 20MB 字节上限,但对整体请求体没有任何字节防线,防线在墙后面。· D3(用失败修复失败):溢出恢复与手动 /compact 均使用 retainTokens = 0(一个字都不留,全量摘要),导致恢复请求和失败请求一样大,自救必然失败。· D4(分类器失明):网关返回的是 HTML 错误页,适配器兜底文案 DeepSeek Messages request failed (413) 无法匹配上下文超限的任何正则,被归类为 INVALID_REQUEST,自动压缩恢复根本不触发。
四、 环境与严重性环境:DSH 0.1.7-alpha.2,Windows 11,Node v24.21.0,deepseek-flash。会话规模:38,400,870 字节 / 40,713 个事件(全程数据未丢失)。本地指标显示“38万/100万,宽裕”,与实际 206 MB 请求体完全脱钩,产生了主动误导。
五、 本地绕过验证(现成补丁)在请求路径上加 8MB 总字节闸门:超限时丢弃超大扩展字段并对图片/超长文本降级,绝不抛错。实测验证:bodyBytes=585413 顺利放行,水位线正常推进,该字段体积从 205.87 MB 缩至 93 KB,死锁自动解除,会话恢复。---详细的全版本(D1-D10)对比、源码追踪(L16行默认值翻转)与完整复现数据,请见下文。
环境
0.1.7-alpha.2webdeepseek-official/deepseek-flashhttps://api.deepseek.com/anthropic现象
每一轮都停在同一个循环里,无法推进:
表面 token 数长期停在 380,109 不再下降,而窗口是 450,000 —— 界面上看不出任何异常。
实测数据(适配器侧抓到的真实 payload,非估算)
dsh_session_log字段baseBytes最后一行是关键证据:消息条数差了 4 倍(208 vs 863),请求体却只差 1.3 MB。说明有一个约 206 MB 的常量在跟着每个请求走,与消息条数无关。
网关响应(原始)
根因链
@deepseek-ai/dsh-session-log-deepseek在 0.1.7-alpha.2 里enabled默认为true(lib/index.jsL16;同机的 0.1.5-rc.3 是false)。dsh_session_log,内容为「上次被接受水位之后的全部会话事件」。session-log-deepseek/delivery-accepted推进,而该事件由accept()写入,dsh-llm-deepseek只在 HTTP 2xx 之后才调用accept()。accept()不执行 → 水位不动 → 下次重发同一份(甚至更大)→ 永久 413。为什么它救不回来(三条共性问题)
以下三条在 0.1.5-rc.3 上同样存在,所以这不是 alpha 回归:
1. 恢复动作会原样重演失败。
dsh-compaction-basic的溢出恢复路径与手动/compact路径都调用selectCompactableRange(session, measurement, 0)。retainTokens = 0的语义是「一个字都不留」,也就是把整个 surface 拿去摘要 —— 恢复请求和刚刚失败的请求一样大,必然同样被拒。用失败的方式去修复失败。2. 错误分类器不认非 JSON 的 provider 报错。
响应体不是 JSON(例如前置网关返回 HTML)时,适配器兜底文案为
DeepSeek Messages request failed (413),该字符串不匹配isContextWindowExceededError()的任何模式,于是被归类为INVALID_REQUEST⇒ 自动压缩恢复根本不触发,连自救机会都没有。分类器假定 provider 一定返回 JSON,没有考虑前置 CDN/网关。3. 请求路径上没有任何「总字节」防线。
压力模型完全建立在 token 上(
contextWindow/thresholdRatio/headroomTokens)。适配器对图片确有 20 MB 的字节上限(maxInlineRequestImageBytes),对整体请求体却一个都没有;且DEFAULT_CONTEXT_WINDOW = 1e6/DEFAULT_MAX_TOKENS = 256e3会把压力压缩的触发线推到约 678k token,远在真实撞墙点之后 —— 防线在墙后面。版本归属(同机两套安装逐条对比)
dsh-session-log-deepseek的enabled默认值truefalseretainTokens = 0恢复路径1e6/256e3默认值false→true的有意翻转,两版 README 措辞正好相反)。这一点值得注意:它把一个 205 MB 的请求体变成了默认行为。修复建议(按重要性)
retainTokens = 0在任何恢复路径上都不安全。bodyBytes—— 只有 token 的指标在存在大扩展字段时是主动误导用户的。--dump-config只打印显式配置、不展开 schema 默认值,导致一个默认开启、能产生 205 MB 行为的功能在任何配置导出里都看不见。本地绕过(供参考)
本机改了三处(3 个文件,适配器部分 +463 / −16 行、13 个 hunk):
dsh-llm-deepseek/lib/index.jsisBodylessStatusdsh-compaction-basic/lib/index.jssummarizeRetainTokens():把retainTokens = 0换成一个既保证「有东西可摘要」、又把摘要输入封顶 131,072 token 的预算enabled: false关掉该遥测关掉遥测之后字段根本不再注册,从源头消除这 205 MB:
第一次请求被放行之后水位正常推进(该字段从 205.87 MB 缩到 93,304 字节),死锁自动解除。会话表面 token 从 380,109 降到 266,393,此后每一轮都正常完成。
补充一:这不是「某个字段太大」,而是「扩展点没有围栏」
本问题可以视为 #4668「Plugins can break the wire protocol」的一个具体实例。
一个可选的遥测贡献有能力把整个 agent 循环打死,原因不是它算错了什么,而是:
accept())挂在请求成功上 —— 于是失败会自我锁死,而不是降级;换句话说:任何未来的插件按同样方式注册一个请求字段,都能复制同一个死法 —— 本次只是第一个被踩到的。
补充二:我们这边的实现(供参考)
先说清四件事:
enabled默认值改回false)。下面这套是通用的字节防线,解决的是「下次换个原因撑爆 body 怎么办」。git apply的.patch随附)。[DSH-REQ]全量请求日志、stderr 落盘 tee),那部分是本机诊断用,不属于建议上游的修复。关键点 1:扩展字段闸门 —— 这一个才是修掉本 bug 的
在
prepareRequestExtensions()里:关键性质:
accept()的调用时机完全不变,只是发出去的载荷变小了。所以第一次请求被放行之后水位正常推进,死锁自己解开 —— 这正是我们实测到的行为(该字段从 205.87 MB 缩到 93,304 字节,之后droppedFields变空)。关键点 2:整包字节阶梯(
fitRequestBody())量的是交给
fetch的那个 JSON 字符串本身;超限就按固定顺序降级,最多 16 轮,永远返回一个 body(降无可降也照发,绝不让回合失败):关键点 3:两个常量与取值理由
关键点 4:分类器 —— 413 按体积处理,结构化 400 不碰
真路径在适配器的
providerError()里,不是dsh-llm的类型定义里(我们一开始改错了文件,见「补充三」):实测行为表:
实测效果(真机)
All reactions