compaction-basic: default headroomTokens (65536) is an absolute value and crushes the compaction trigger on small context windows
#8079
Replies: 3 comments 1 reply
|
Cross-reference: #8088 reports the same symptom independently — "[BUG] Compaction happens every few steps for no reason, context is only around 50% of maximum" — on a different machine and configuration. Adding it here as a second data point, since two independent reproductions plus the root cause should be more useful to the maintainers than either alone. Also related: #76 ( |
|
I just moved to dsh-v0.2.0-rc.2 and this bit me too. Degraded performance due to frequent compactions and resulting loss of context detail causing an excessive increase in overall task time. Not usable as-is with a 180K, 128K, or smaller I think the better fix is not just documenting that headroomTokens has to be tuned on smaller routes, but changing the trigger so it reserves the headroom the compaction request actually needs. Right now the default uses a flat 65536, which makes sense as a conservative cloud default but is disproportionately expensive on 128K to 200K local models. A more accurate trigger would be based on the auxiliary summary call itself: drop the retained tail from the prompt, add the compaction instruction, reserve the summary call’s own maxTokens budget, and maybe keep a small fixed safety margin. In other words, make the threshold reflect “can the summary request still fit?” rather than “always subtract 64K no matter the route.” Related to that, the summary call should probably avoid replaying prompt baggage that is only needed for the main tool-using request path. If the summarizer does not need the full tool schema/tool-call prompt context to produce a good history summary, excluding that from the compaction request would reduce the required reserve further and make the trigger behave much more like W - O on smaller local windows. That seems like a better long-term direction than keeping a large absolute headroom and asking users to discover and tune around it per route. 中文翻译: 我刚升级到 dsh‑v0.2.0‑rc.2,也同样被这个问题坑到了。由于频繁压缩导致性能下降,加上上下文细节不断丢失,使整体任务时间大幅增加。在 180K、128K 或更小的 W 最大上下文、以及 1/3 O 输出上限的配置下,这个行为基本无法正常使用。 我认为更好的修复方式不只是记录在案、告诉用户在较小的路线上必须调节 headroomTokens,而是直接修改触发逻辑,让它实际为压缩请求预留所需的余量。现在的默认值是固定的 65536,这在云端的大窗口上作为保守默认是合理的,但在 128K 到 200K 的本地模型上成本却高得不成比例。更准确的触发方式应该基于辅助的 summary 调用本身:从提示中去掉保留的尾部,加上压缩指令,预留 summary 调用自身的 maxTokens 配额,并可能保留一个小的固定安全边界。换句话说,让阈值反映“summary 请求还能塞得进去吗?”而不是“无论路线如何都固定减去 64K”。 与此相关的是,summary 调用应该避免重放那些只在主工具调用路径中才需要的提示包袱。如果总结器在生成历史摘要时并不需要完整的工具 schema / 工具调用提示上下文,那么在压缩请求中排除这些内容可以进一步减少所需的预留空间,使触发行为更像在较小本地窗口上的 W − O。相比维持一个大的绝对 headroom 并让用户在每条路线中自行摸索和调参,这似乎是更好的长期方向。 |
|
Thanks for the report. |
Uh oh!
There was an error while loading. Please reload this page.
Summary
In 0.1.7 the compaction trigger in
@deepseek-ai/dsh-compaction-basicbecame a two-term minimum:headroomTokensdefaults to 65536 and is an absolute number that does not scale with thewindow. That default is harmless for a 1M-token cloud route but very costly for a local
route: the second term wins, so
thresholdRatio: 0.8stops having any effect at all.Environment
0.1.7-alpha.2, Windows + WSL2, npm-global installcontextWindow: 150000, requestmaxTokens: 32000(reserved completion)contextWindow: 126976, requestmaxTokens: 32000Impact
Once the harness's own fixed overhead is counted (
@deepseek-ai/dsh-token-meterfolds thetool schemas into the measurement — ~18K tokens with a plugin-heavy tool set here), a 150K
route leaves only ~34K tokens of actual conversation before compaction fires. The same
65536 costs a 1M route only ~6.6 percentage points.
At 40K the pressure budget goes negative,
resolveCompactSpecthrowsTargetPressureConfigError, theagent/pre-steplistener catches it and warns once — sothere is no pressure trigger on that route at all, which is the opposite failure mode.
Measured evidence (session logs)
Decoding
~/.dsh/sessions/**/session.v4.jsonl.zstd(one zstd frame per append) and countingcompaction/*againststep/*:compaction/startcompaction/prunestep/startOne 49-step session compacted 32 times — roughly once every 1.5 steps.
Secondary effect: the pruner then throws away what was just read
Every
compaction/pruneevent in those runs hasshadowedRange.start == end, i.e. a singleevent is pruned, and
shadowedTokenCountfor them is 2,246–10,145 tokens (whole-file reads).Because
compactIfNeededprunes before it re-measures, an entry can lose its freshly readcontent without producing a summary:
That produced a self-sustaining "read file → get pruned → read the same file again" loop
(same file read 20+ times in one session).
Suggested fix
Any one of these would address it:
Math.min(65536, Math.floor(contextWindow * 0.1)),or make the second term
messageBudgetTokens * (1 - thresholdRatio)so the ratio term keepsmeaning. (This is what we ended up configuring locally:
headroomTokens: 8192, which puts a150K route back at
min(120000, 150000-32000-8192) = 109,808≈ 73%.)headroomTokensto 0 and letthresholdRatiobe the only knob, documenting theabsolute-headroom option for users who want it.
maxTokensrather than a fixed 65536.Also worth documenting on the
compaction-basicreference page: the trigger ismin(W × ratio, W − reserved − headroom), the fixed tool-schema overhead is counted, and on a128K route the usable conversation budget is roughly
W − reserved − headroom − tools.Workaround
Through 0.1.7 the value has to be set in the preset that defines the live composition
(
dsh-web-app/presets/standard.patch.yml,ptc.patch.yml,cordis.patch.yml):Note that a profile
cordis.patch.ymloverride targeting- id: compaction-basichas noeffect — the dsh-base top-level row is
disabled: trueunder a preset, so the override landson the disabled row while the preset's own copy stays at the default. Worth a note in the
preset/profile patch documentation.
All reactions