-
Notifications
You must be signed in to change notification settings - Fork 5
Dynamic Compaction Decision: When to Compact and How Much to Keep
Most AI coding agents compact context with a straightforward rule: when the token count exceeds a threshold, summarize and keep the last N messages. That is simple, but it misses three costs that matter in long sessions:
- Compaction itself costs money: the summary call still consumes input and output tokens.
- Compaction changes the cache prefix: the new summary plus retained messages introduce a one-time cache miss cost.
- Long context can hurt quality: keeping everything may preserve details, but long prompts can make the model miss relevant information in the middle.
bash-agent uses a cache-aware DP economics model. On each compaction check it answers two questions:
- Is compaction worth doing?
- If yes, how many recent conversation lines should be retained?
If every candidate has non-positive net benefit, compaction is skipped. When compaction is forced, for example near the context limit or by PlanClear / PlanConfirm, the same turn-aligned retention policy is used as a fallback.
The model enumerates k, the number of recent conversation JSONL lines to retain. For every candidate k it estimates:
-
$K$ : tokens in the retained recent$k$ lines. -
$H$ : old history tokens that would be replaced by the summary. -
$T$ : total estimated tokens in the conversation file.
The model chooses the highest positive net benefit and then moves the cut point back to a user-message boundary so it never truncates the middle of an assistant/tool segment.
The model computes five terms:
Where:
-
$E$ : expected remaining user-input turns. By default it is estimated fromDP_BASELINE_E - current_turn_countwith a floor;DP_E_FIXEDcan force a fixed value. -
$L$ : average LLM calls per user input.DP_L=0auto-estimates it fromagent_request_count / current_turn_count; otherwise it uses the configured value. -
$R = E \times L$ : expected remaining LLM calls. -
$\text{avg}$ : average input tokens per LLM request, estimated from cumulative stats, defaulting to 4000. -
$r_t = \max(r^{c+1}, 0.37)$ : cumulative retention after this compaction, with a floor to avoid runaway penalties. -
$M$ : max context tokens fromMAX_CONTEXT_TOKENS, default 200000. -
$Q$ :DP_QUALITY_PENALTY, default 0.2.
After compaction, later requests carry
After compaction, the prefix becomes “new summary + retained recent messages”. That prefix is new on the next request, so the model accounts for the difference between full input price and cached input price.
The summary call reuses the normal request prefix. The fixed prefix
Summaries lose detail. The implementation estimates the loss using cumulative retention after this compaction,
This is the newer positive term. It models the fact that very long contexts degrade answer quality. Compacting shortens the working context from
| Parameter | Default | Description |
|---|---|---|
DP_P_INPUT |
3.0 |
Uncached input price, $/MTok |
DP_P_CACHE |
0.30 |
Cache-hit input price, $/MTok |
DP_P_OUT |
15.0 |
Output price, $/MTok |
DP_V |
5000 |
Fixed prefix tokens: system prompt, tools, old summary, etc. |
DP_S |
500 |
Estimated summary output tokens |
DP_L |
0 |
Average LLM calls per user input; 0 means auto-estimate |
DP_BASELINE_E |
8 |
Baseline expected remaining user-input turns |
DP_E_FIXED |
0 |
Fixed E override; values greater than 0 skip dynamic estimation |
DP_R |
0.8 |
Single-summary information retention rate |
DP_BETA |
0.03 |
Information distortion penalty coefficient |
DP_QUALITY_PENALTY |
0.2 |
Long-context quality decay penalty coefficient |
DP_MIN_KEEP_RATIO |
0.12 |
Minimum ratio of message lines to retain |
MAX_CONTEXT_TOKENS |
200000 |
Max context limit, used for forced compaction and quality-term normalization |
- Per-line tokens are approximated as
(byte_length + 3) / 4 + 1. -
min_keepkeeps at least 3 lines and respectsDP_MIN_KEEP_RATIO. - The model evaluates every
k ∈ [min_keep, NR]. - Automatic compaction only happens when
best_benefit > 0. - The chosen cut point is moved back to a user-message boundary.
- Context above 90% of the limit,
plan_clear, andplan_confirmcan force a fallback compaction.
The bash-agent compaction model is not simply “compact when over threshold”. It weighs:
- cache-priced savings from dropping old history;
- one-time cache invalidation from the new summary prefix;
- the summary request cost itself;
- information loss from repeated summarization;
- quality gains from shortening an overlong context.
This avoids compacting repeatedly near the boundary while still compacting more aggressively when the working context itself is likely hurting answer quality.
Related article: Cache-Aligned Summarization
Source code:compact_dp.awk,agent.sh