Replies: 1 comment
Workaround (profile-local plugin) / 可用的临时方案(配置本地插件)Until the header-inheritance fix lands, the cache-breaking half of this bug can be worked around without forking any package: a small profile-local Cordis plugin that listens on the global 在修复合入之前,本 bug 中「破坏缓存」的那一半可以不动任何包源码、用一个配置本地的小型 Cordis 插件绕过:监听全局 How to install / 安装方法 (tested on
const name = "compaction-reasoning";
const inject = ["sessions"];
function apply(ctx) {
ctx.on("llm/stream", (options, next) => {
if (options?.purpose !== "compaction") return next();
if (options.reasoningEffort !== void 0 && options.temperature !== void 0) return next();
if (options.sessionId === void 0) return next();
const session = ctx.sessions.get(options.sessionId);
if (session === void 0) return next();
const header = session.requestHeader();
const config = header?.config;
if (config === void 0) return next();
if (options.reasoningEffort === void 0 && config.reasoningEffort !== void 0 && header?.adapterDefaults?.reasoningEffort !== true) {
options.reasoningEffort = config.reasoningEffort;
}
if (options.temperature === void 0 && config.temperature !== void 0) {
options.temperature = config.temperature;
}
return next();
}, { global: true, prepend: true });
}
export { apply, inject, name };
- insert:
- id: compaction-reasoning
name: './plugins/compaction-reasoning/lib/index.js'
Scope / 适用范围: the listener is preset-independent (registered globally), so it covers every preset and both automatic and 该监听器全局注册、与预设无关,自动压缩与 Long-term / 长期方案: the proper fix is still the one proposed above — carry the effective 长期方案仍是楼上建议的做法:把路由请求头中生效的 |
Uh oh!
There was an error while loading. Please reload this page.
English
Summary
dsh-compaction-basiccan invalidate the provider's warm KV prefix when the active session uses a per-request reasoning effort different from the provider default. The compaction summarizer carriesproviderandmodel, but notreasoningEffort. For providers whose chat template encodes reasoning effort, the auxiliary compaction prompt diverges near the beginning and the intended prefix-cache reuse is lost.A second problem makes this especially painful for local models:
dsh-llm-pi-aicancels a stream after 300,000 ms without a parsed model chunk. A local server can still be actively prefilling a large prompt, but transport keepalives or prompt-progress events do not appear to pulse this watchdog. The Harness cancels healthy work before the server's longer request timeout.Environment
0.1.7-rc.1, Web profiledsh-llm-pi-ai, OpenAI-compatible local endpointxhighmediumstreamIdleTimeoutMs: default 300,000 msReproduction and evidence
The normal conversation used
xhigh. Automatic compaction started around 149K input tokens. Four compaction attempts were observed:The successful request took about 342 seconds in total, but it survived because output began after 152.4 seconds and subsequent chunks reset the idle watchdog. The cancelled attempts appear to have progressively populated a cache for the compaction prompt, eventually allowing attempt 4 to reach its first token in time.
A normal request immediately after compaction showed the same independent watchdog behavior: a 57,413-token cold request was cancelled after 300 seconds; its retry reused 8,384 tokens and succeeded with TTFT 279.4 seconds.
Code path
The current
summarizeWithLlm()chooses a target containing onlyproviderandmodel, and itsGenerateOptionsomitsreasoningEffort:https://github.com/deepseek-ai/deepseek-harness/blob/master/packages/compaction/compaction-basic/src/summarizer.ts
The Pi AI adapter then resolves:
The repository's own compaction design note says the replayed request should reproduce the warm prefix verbatim. Omitting a template-affecting request option violates that invariant:
https://github.com/deepseek-ai/deepseek-harness/blob/master/.agents/notes/archived/bug-fix/2026-07-21-compaction-summary-prefix-cache-reuse.md
For this model,
mediumandxhighproduce different system-template text almost immediately, so the whole prior prefix becomes unusable.Expected behavior
Suggested fixes
reasoningEffortfrom the routed request header into the summarization call. More generally, preserve every request option that can alter prompt rendering.purpose: "compaction", since these requests are predictably the largest.Workaround
Users can set a larger
streamIdleTimeoutMsin their provider profile, but that only prevents premature cancellation. It does not repair the cache-breaking reasoning mismatch.中文
摘要
当当前会话使用的逐请求推理强度与 provider 默认值不同时,
dsh-compaction-basic可能会使 provider 已预热的 KV 前缀缓存失效。压缩摘要请求只继承了provider和model,没有继承reasoningEffort。对于把推理强度写入聊天模板的 provider,辅助压缩请求会从提示词开头附近就发生变化,因此无法实现预期的前缀缓存复用。第二个问题会让本地模型尤其容易失败:如果连续 300,000 ms 没有收到已解析的模型数据块,
dsh-llm-pi-ai就会取消流。此时本地服务器可能仍在正常处理一个很长的输入,但传输 keepalive 或提示词处理进度似乎不会重置该 watchdog。因此 Harness 会在服务器自身更长的请求超时之前取消仍然健康的工作。环境
0.1.7-rc.1,Web profiledsh-llm-pi-ai,本地 OpenAI 兼容端点xhighmediumstreamIdleTimeoutMs:默认 300,000 ms复现与证据
正常会话使用
xhigh。输入约达到 149K token 时触发自动压缩。观察到四次压缩请求:成功的第四次请求总共约耗时 342 秒,但因为 152.4 秒后已经开始输出,后续数据块不断重置 idle watchdog,所以能够完成。前三次被取消的请求似乎逐步建立了压缩请求自己的缓存,最终使第四次请求能及时产生首 token。
压缩成功后的普通请求也独立复现了 watchdog 问题:一个 57,413-token 的冷请求在 300 秒后被取消;重试复用了 8,384 token,并以 279.4 秒 TTFT 成功。
代码路径
当前
summarizeWithLlm()选择的 target 只包含provider和model,构造的GenerateOptions没有reasoningEffort:https://github.com/deepseek-ai/deepseek-harness/blob/master/packages/compaction/compaction-basic/src/summarizer.ts
随后 Pi AI adapter 使用以下逻辑:
仓库自己的压缩设计说明强调,重放请求应逐字复现已预热的前缀。遗漏会改变模板的请求选项违反了这一约束:
https://github.com/deepseek-ai/deepseek-harness/blob/master/.agents/notes/archived/bug-fix/2026-07-21-compaction-summary-prefix-cache-reuse.md
对于本次使用的模型,
medium和xhigh几乎从 system 模板开头就不同,因此之前的整个前缀缓存无法使用。预期行为
建议修复
reasoningEffort传给摘要请求。更一般地说,应保留所有会改变提示词渲染的请求选项。purpose: "compaction"设置单独或更长的默认 idle 限制,因为压缩请求通常就是最大的请求。临时解决方法
用户可以在 provider profile 中增大
streamIdleTimeoutMs,这只能避免过早取消,不能修复由 reasoning 不一致导致的缓存失效。All reactions