Prefix-cache-friendly compaction: rewrite-the-prefix vs append-only reconstruction #5638
Replies: 1 comment
|
Disclosure: I build Grunz, a hosted agent on open-weight models, and we went through roughly this redesign. One constraint I'd add to direction A, because moving everything mutable to the tail has a cost that's easy to miss. Split the durable block by how often it changes, not all-or-nothing. The table in section 3 mixes two kinds of content:
If both go to the tail, the cache problem is solved, but the contract then sits after a growing history and competes for attention with the most recent tool output. That is the text we saw agents lose first, well before any summarizer touched it. We ended up with the contract as a frozen block right after the system prompt (rewritten only when the user actually changes the task, so it costs one miss per change) and the ledger appended at the tail as deltas. That gives you pi's frozen prefix and keeps the goal where the model weighs it most. Budget the frozen block before computing the trigger. Our first version re-injected the goal without reserving space for it. On a small window that moved the trigger earlier and compaction started firing in a loop, because the part it wasn't allowed to touch was part of the overflow. On the 32K trigger in the example config: that happens to match the window many hosted OpenAI-compatible endpoints actually serve, whatever the model card says, so it's a safe default for them. On a real 128K-1M endpoint, though, it's what makes the rewrite frequent, as you note. Deriving the trigger from the resolved served window (with its source recorded) rather than a fixed token count would help with both the cache and the goal-retention side. |
Uh oh!
There was an error while loading. Please reload this page.
This is a design discussion, not a bug report. Compaction, durable-context injection, and a static system prompt are all working as designed. The question is whether that design is fighting prompt prefix cache (the provider-side reuse of an identical prompt prefix across consecutive model calls; often billed at ~10% of uncached input).
DeerFlow already treats prefix cache as a first-class constraint for the system prompt. Compaction currently rewrites the conversation prefix, which can throw away a cache a long tool loop had been warming. This post compares that path with pi's append-only reconstruction and asks whether we want to move in that direction.
Related (different scope, please do not merge this into them):
Why prefix cache cares about byte-identical prefixes
Providers (OpenAI, Anthropic, and many OpenAI-compatible backends) cache the leading bytes of a prompt. The next call only reuses that cache if the new prompt starts with the same bytes. Anything that mutates an earlier message — a new summary at the front, a live ledger inserted after the system prompt, a rewritten tool result in the middle — invalidates the cache from that point onward.
A typical agent tool loop is the best case for this:
[system][user]→ cache write[system][user][ai][tool]→ hit on[system][user], write the rest[system][user][ai][tool][ai][tool]→ hit on the previous prefixEach step only pays for the newly appended tokens. One rewrite of the prefix turns the next call into a full miss.
DeerFlow already encodes this for the system prompt:
DynamicContextMiddlewareinjects date/memory as a frozen reminder so the system prompt stays static (agent.pyaround the "prefix-cache reuse" comment;dynamic_context_middleware.pymodule docstring).SystemMessageCoalescingMiddlewarepasses through with zero mutation when there is nothing to merge, "preserving prefix-cache hits".The gap is the conversation prefix after the system prompt, which compaction and several request-time injectors rewrite.
What DeerFlow does today
Shipped
config.example.yaml: summarizationenabled: true, trigger at 32k tokens, keep the last 10 messages.1. Compaction replaces the whole message list
DeerFlowSummarizationMiddleware._maybe_summarize/_amaybe_summarize(summarization_middleware.py):Older turns are deleted from checkpoint state. The next model call is no longer "the previous prompt plus one new message". It is a new prompt: remaining tail + a newly generated
summary_text._preserve_required_contextcan also reorder: system messages, tagged dynamic-context reminders, and the current user request are rescued out of the summarize window and placed with the kept tail. The kept list is therefore not a stable suffix of the original list.This is expected at the moment of compaction (you cannot drop history without changing the prefix once). The issue is what happens after that, and how often it happens.
2. The keep window is a sliding rewrite, not a frozen tail
Default keep is 10 messages (example config) / 20 messages (schema default). The next compaction summarizes part of that previously kept tail, writes a new
summary_text, and keeps a new last-N. The post-compact prefix never gets a chance to sit still across many turns.On a 128k–1M model, firing at 32k (RFC #4346 proposal B) makes this rewrite frequent.
3. Live durable context is inserted at the front of every request
DurableContextMiddleware._inject(durable_context_middleware.py) does not persist the projection. On everywrap_model_callit rebuilds:via
insert_after_leading_system_messages. The data block mixes:summary_textdelegationsskill_contexttask_notes/task_historyAny of those changing rewrites the message that sits between the system prompt and the history. From the provider's point of view the prefix broke, so the entire kept conversation misses cache — even when summarization did not run on that turn.
SystemMessageCoalescingMiddlewarethen folds the extra authoritySystemMessageinto the leading system block. A change there mutates the system prefix itself.4. Other front-insert / in-prefix rewrites
ToolReceiptMiddlewareinserts a receipt ledger with the sameinsert_after_leading_system_messageshelper. As receipts accumulate, that front-inserted HumanMessage grows.ToolOutputBudgetMiddleware._patch_model_messages/elide_superseded_write_payloadsrewrite historical tool payloads in the request copy only. The elision policy is monotonic (once gone, it stays gone), so that path only misses cache once per elided write, which is the right shape. It is listed here for completeness, not as the main problem.5. Compaction can fire mid tool-loop
Summarization runs in
before_model. During a long tool loop the hottest cache is the growing append-only prefix of the current turn. A compact between tool calls discards that prefix just when reuse would have been cheapest.What pi does instead
Reference: pi compaction docs, session format,
packages/coding-agent/src/core/compaction/compaction.ts.Pi also puts a summary at the front of the model prompt after a compact. The important difference is that the post-compact prefix then stops moving until the next compact.
Append-only session log. Compaction appends a
CompactionEntry{ summary, firstKeptEntryId, tokensBefore, details }. Original messages are not deleted or rewritten in the log.Identity-stable reconstruction.
buildContextEntries()uses the latest compaction:firstKeptEntryIdup to the compaction entry (same objects, same bytes)After one compact,
[summary][kept original messages]is frozen. Later turns only append. The next N model calls can hit that entire prefix.Large frozen tail, late trigger. Defaults:
keepRecentTokens = 20_000, compact whencontextTokens > contextWindow - reserveTokens(reserveTokens = 16_384). The post-compact prefix is large enough to be worth caching, and compaction is rare enough that the cache can warm.Summarization call does not pollute cache.
completeSummarizationsetscacheRetention: "none"and uses a fresh routing session ID, because that one-off prompt will not be reused.No live-mutating front blob. File tracking is stored in compaction
detailsand baked into that snapshot. Later file ops do not rewrite the summary message sitting at the front of the prompt.Toy picture:
Directions (questions, not a PR)
Not proposing a patch yet. Options, independently mixable:
A. Placement of mutating context. Stop putting a live-updating durable-context / receipt block between the system prompt and history. Keep a frozen snapshot in the prefix (the last compaction summary) and append ledger/skill/receipt deltas at the tail — or inject them only when they actually changed, as a new message, never by rewriting an earlier one.
Trade-off:
insert_after_leading_system_messagesexists so injected context is not read as the latest user/tool turn and so it cannot precede the system prompt. Any move-to-tail design has to preserve that.B. Compaction identity (
firstKeptEntryId). StopREMOVE_ALL_MESSAGES+ last-N rewrite. Treat the kept tail as the original messages from a cut point, and append a compaction record. Subsequent turns append after that cut. Checkpoint / UI / event-store implications are the hard part (history display currently benefits from event-store recovery of removed messages).C. When to compact. Compact near remaining window headroom (pi's
reserveTokens) rather than an absolute 32k / last-10-messages. Overlaps RFC #4346 proposal B; listed here only because trigger frequency is a cache-rate knob.D. Compact at turn boundaries. Skip automatic compaction between tool calls of the same user turn unless the next call would overflow. Lets the hottest append-only prefix survive.
E. Summarization request hygiene. Disable prompt-cache writes on the one-off summary LLM call (
cacheRetention: "none"/ equivalent). Small, independent of A–D.What this is not
summary_text/ ledger / skills) that pi does not have to express the same way.Asks
prompt_cache_key/ session affinity)?Happy to follow up with a more concrete design (or a small probe: log
cache_read_tokensbefore/after one compact on a long tool loop) once direction is clear.中文摘要 / Chinese summary
这不是 bug,是设计取舍。模型的前缀缓存(prompt prefix cache)要求连续两次请求的开头字节完全一样,命中时输入会便宜很多。DeerFlow 已经把系统提示词做成静态的,就是为了吃到这次缓存;但当前的会话压缩会把对话前缀整段换掉,长工具循环里刚攒起来的缓存会一次打空。
现在怎么做: 压缩触发后用
REMOVE_ALL_MESSAGES删掉旧消息,只留下最近 N 条(示例配置是 10 条)和新生成的摘要。摘要、委派台账、技能列表还会在每一次模型请求里,插到系统提示词和对话历史中间。台账/技能一变,这块内容就变,后面整段历史都无法命中缓存——即使这一轮并没有再次压缩。压缩还可能发生在同一轮工具调用之间,那正好是缓存最值钱的时候。pi 的做法: 会话只追加、不改写。压缩时追加一条带
firstKeptEntryId的记录;发给模型的是「摘要快照 + 从该 id 起的原始消息 + 压缩之后新追加的消息」。压缩一次之后,这段前缀会冻住,直到下一次压缩。默认保留约 2 万 token、接近窗口上限才压。摘要那一次请求还会关掉 cache 写入,避免一次性 prompt 污染缓存。想请维护者拍板:压缩后的前缀缓存是否作为目标?上面 A–E(注入位置 / 按 id 保留尾巴 / 触发时机 / 不要在工具循环中途压 / 摘要请求关 cache 写入)哪些可以做、哪些明确不做。
All reactions