You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Compaction summarization drops the session's reasoningEffort, destroying the prompt cache on models that render effort into the system prompt
Package:@deepseek-ai/dsh-compaction-basic (observed in dsh 0.1.5-rc.2) Provider:@deepseek-ai/dsh-llm-pi-ai, api: openai-completions against a local llama.cpp server Model: Qwen3.8-27B (GGUF, local), 240k→150k context window
Summary
summarizeWithLlm() builds the compaction request from the session's own request
header, but copies only provider and model from it — not reasoningEffort.
The request therefore reaches the adapter with no effort, and dsh-llm's resolveCallWithInfo() fills in the provider default
(llm-pi-ai.providers.<route>.reasoning).
For any session whose effort differs from that provider default, the
compaction request's rendered prompt differs from the conversation it is
summarizing. On a provider that caches prompt prefixes this silently discards
the whole cache; on my setup (a single local GPU) that turned a ~10-second
operation into a 16-to-27-minute one.
The function is explicitly written to be cache-friendly — it replays the
conversation prefix and appends the compaction instruction as the final user
message "so the provider's warm prefix cache is reused" — so I believe the
omission is an oversight rather than intent.
Why the effort ends up in the prompt
This is not specific to my deployment, but it is what makes the bug expensive.
Qwen3.8's chat template (shipped inside the GGUF) renders the reasoning effort
as the first line of the system prompt:
<|im_start|>system
Reasoning effort is set to xhigh. Please think carefully through the task, ...
# Tools
...
With medium that line is omitted entirely. So the prompts diverge at token 3,
and every cached token after it is worthless. Any model family that turns reasoning_effort into prompt text (rather than a pure sampling parameter) has
the same exposure.
Reproduction
Configure one pi-ai provider route with a default effort, e.g.
Run a session on that route with a different effort — for example a
subagent started with reasoning_effort: low
(subagent-model-selection.enabled: true).
Let the session cross the compaction threshold.
Inspect the request bodies on the wire. The conversation's own requests carry "reasoning_effort": "low"; the compaction request carries "reasoning_effort": "xhigh".
I verified both directions with a stub server that logs the reasoning_effort
of every request, and in production against llama.cpp.
Observed cost (real runs, llama.cpp server logs)
case
session effort
compaction effort
prompt reused
main chat, provider default medium
xhigh
medium
0 of 114 308 tokens (27 min re-prefill)
subagent, provider default xhigh
low
xhigh
8 775 of 91 237 tokens (16.5 min re-prefill)
after the patch below
medium
medium
85 154 of 86 582 tokens (98 %)
llama.cpp reports the shared prefix directly:
slot operator(): id 1 | task 66275 | new prompt, n_ctx_slot = 163840, task.n_tokens = 91237
slot operator(): id 1 | task 66275 | checking checkpoint with [87240, 87240] against 8778...
slot operator(): id 1 | task 66275 | cached n_tokens = 8775, memory_seq_rm [8775, end)
against 8778 is the common prefix — i.e. everything past the shared tool
preamble was recomputed.
Why a provider default cannot work around it
There is no single value that is correct for both sides: the main chat runs at
one effort and its subagents deliberately run at another (cheaper reasoning for
well-specified execution work). Whichever value the provider default carries,
the other side pays a full re-prefill at every compaction.
Suggested fix
Inherit the session's effort when the summarization runs on the same route. latest already holds it:
// lib/index.js, summarizeWithLlm()constlatest=agent.session.requestHeader()?.config;
...
constoptions={provider: target.provider,model: target.model,
messages,
...input.tools===void0 ? {} : {tools: [...input.tools]},// inherit the session's effort when summarizing on the same route,// so the replayed prefix renders identically and stays cache-hot
...latest?.reasoningEffort===void0||latest.provider!==target.provider||latest.model!==target.model ? {} : {reasoningEffort: latest.reasoningEffort},maxTokens: config.maxTokens,sessionId: agent.session.id,purpose: "compaction",
...signal===void0 ? {} : { signal },};
The route guard matters: when summarizationProvider/summarizationModel point
at a different model, that model may not accept the session's effort id, and dsh-llm would reject the call with UNSUPPORTED_REASONING_EFFORT.
An explicit BasicCompactionConfig.reasoningEffort (defaulting to "inherit")
would work as well, and would let deployments force a cheap effort for
summaries where the cache does not matter.
Related, same root cause
dsh-session-title-first-prompt-llm also issues its request without an effort
and picks up the provider default. That one is harmless in practice — it is a
separate short conversation with its own cache entry — but the same inheritance
question applies.
Environment
dsh 0.1.5-rc.2, --profile web, Node 24.21.0, Linux
custom agent preset copied from standard, with compaction-basic.maxTokens: 40960 and tool-result-pruner disabled
backend: llama.cpp (build 10460), Vulkan, single GPU, 2 slots, prompt caching
and context checkpoints enabled
Note on compaction-basic.maxTokens
Separately, the default maxTokens: 8192 is too small once the summarization
inherits a reasoning-heavy effort: the model spends the budget thinking and the
run fails with summarization truncated at the token cap (incomplete checkpoint). In one of my runs it failed three times in a row, the context kept
growing unchecked, and the subagent ended on max-tokens after 3 h 15 of work.
Raising it to 40960 fixed that, but a default that cannot fit a single
reasoning-model summary is an easy trap — perhaps derive it from the model's
reasoning budget, or at least surface the repeated truncation more loudly.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Compaction summarization drops the session's
reasoningEffort, destroying the prompt cache on models that render effort into the system promptPackage:
@deepseek-ai/dsh-compaction-basic(observed indsh0.1.5-rc.2)Provider:
@deepseek-ai/dsh-llm-pi-ai,api: openai-completionsagainst a local llama.cpp serverModel: Qwen3.8-27B (GGUF, local), 240k→150k context window
Summary
summarizeWithLlm()builds the compaction request from the session's own requestheader, but copies only
providerandmodelfrom it — notreasoningEffort.The request therefore reaches the adapter with no effort, and
dsh-llm'sresolveCallWithInfo()fills in the provider default(
llm-pi-ai.providers.<route>.reasoning).For any session whose effort differs from that provider default, the
compaction request's rendered prompt differs from the conversation it is
summarizing. On a provider that caches prompt prefixes this silently discards
the whole cache; on my setup (a single local GPU) that turned a ~10-second
operation into a 16-to-27-minute one.
The function is explicitly written to be cache-friendly — it replays the
conversation prefix and appends the compaction instruction as the final user
message "so the provider's warm prefix cache is reused" — so I believe the
omission is an oversight rather than intent.
Why the effort ends up in the prompt
This is not specific to my deployment, but it is what makes the bug expensive.
Qwen3.8's chat template (shipped inside the GGUF) renders the reasoning effort
as the first line of the system prompt:
With
mediumthat line is omitted entirely. So the prompts diverge at token 3,and every cached token after it is worthless. Any model family that turns
reasoning_effortinto prompt text (rather than a pure sampling parameter) hasthe same exposure.
Reproduction
Configure one pi-ai provider route with a default effort, e.g.
Run a session on that route with a different effort — for example a
subagent started with
reasoning_effort: low(
subagent-model-selection.enabled: true).Let the session cross the compaction threshold.
Inspect the request bodies on the wire. The conversation's own requests carry
"reasoning_effort": "low"; the compaction request carries"reasoning_effort": "xhigh".I verified both directions with a stub server that logs the
reasoning_effortof every request, and in production against llama.cpp.
Observed cost (real runs, llama.cpp server logs)
mediumxhighllama.cpp reports the shared prefix directly:
against 8778is the common prefix — i.e. everything past the shared toolpreamble was recomputed.
Why a provider default cannot work around it
There is no single value that is correct for both sides: the main chat runs at
one effort and its subagents deliberately run at another (cheaper reasoning for
well-specified execution work). Whichever value the provider default carries,
the other side pays a full re-prefill at every compaction.
Suggested fix
Inherit the session's effort when the summarization runs on the same route.
latestalready holds it:The route guard matters: when
summarizationProvider/summarizationModelpointat a different model, that model may not accept the session's effort id, and
dsh-llmwould reject the call withUNSUPPORTED_REASONING_EFFORT.An explicit
BasicCompactionConfig.reasoningEffort(defaulting to "inherit")would work as well, and would let deployments force a cheap effort for
summaries where the cache does not matter.
Related, same root cause
dsh-session-title-first-prompt-llmalso issues its request without an effortand picks up the provider default. That one is harmless in practice — it is a
separate short conversation with its own cache entry — but the same inheritance
question applies.
Environment
dsh0.1.5-rc.2,--profile web, Node 24.21.0, Linuxstandard, withcompaction-basic.maxTokens: 40960andtool-result-prunerdisabledand context checkpoints enabled
Note on
compaction-basic.maxTokensSeparately, the default
maxTokens: 8192is too small once the summarizationinherits a reasoning-heavy effort: the model spends the budget thinking and the
run fails with
summarization truncated at the token cap (incomplete checkpoint). In one of my runs it failed three times in a row, the context keptgrowing unchecked, and the subagent ended on
max-tokensafter 3 h 15 of work.Raising it to 40960 fixed that, but a default that cannot fit a single
reasoning-model summary is an easy trap — perhaps derive it from the model's
reasoning budget, or at least surface the repeated truncation more loudly.
All reactions