Skip to content

Dynamic Compaction Decision: When to Compact and How Much to Keep

lloydzhou edited this page May 29, 2026 · 4 revisions

1. Beyond Simple Thresholds

Most AI coding agents compact context with a straightforward rule: when the token count exceeds a threshold, summarize and keep the last N messages. That is simple, but it misses three costs that matter in long sessions:

  • Compaction itself costs money: the summary call still consumes input and output tokens.
  • Compaction changes the cache prefix: the new summary plus retained messages introduce a one-time cache miss cost.
  • Long context can hurt quality: keeping everything may preserve details, but long prompts can make the model miss relevant information in the middle.

bash-agent uses a cache-aware DP economics model. On each compaction check it answers two questions:

  1. Is compaction worth doing?
  2. If yes, how many recent conversation lines should be retained?

If every candidate has non-positive net benefit, compaction is skipped. When compaction is forced, for example near the context limit or by PlanClear / PlanConfirm, the same turn-aligned retention policy is used as a fallback.

2. Decision Variable

The model enumerates k, the number of recent conversation JSONL lines to retain. For every candidate k it estimates:

  • $K$: tokens in the retained recent $k$ lines.
  • $H$: old history tokens that would be replaced by the summary.
  • $T$: total estimated tokens in the conversation file.

The model chooses the highest positive net benefit and then moves the cut point back to a user-message boundary so it never truncates the middle of an assistant/tool segment.

3. Current 5-Term Net Benefit Formula

The model computes five terms:

$$ \begin{aligned} \text{NetBenefit}(k) &= \underbrace{\frac{(R - 1) \cdot P_{\text{cache}} \cdot H}{10^6}}_{①;\text{future savings}} \\ &\quad -\underbrace{\frac{(S + K) \cdot (P_{\text{input}} - P_{\text{cache}})}{10^6}}_{②;\text{cache miss}} \\ &\quad -\underbrace{\frac{P_{\text{cache}}(V + H) + P_{\text{input}} \cdot L_{\text{instr}} + P_{\text{out}} \cdot S}{10^6}}_{③;\text{compaction request cost}} \\ &\quad -\underbrace{\frac{\beta \cdot (1 - r_t) \cdot R \cdot \text{avg} \cdot P_{\text{input}}}{10^6}}_{④;\text{information distortion penalty}} \\ &\quad +\underbrace{Q \cdot P_{\text{input}} \cdot \frac{(V + T)^2 - (V + K)^2}{M \cdot 10^6}}_{⑤;\text{quality improvement savings}} \end{aligned} $$

Where:

  • $E$: expected remaining user-input turns. By default it is estimated from DP_BASELINE_E - current_turn_count with a floor; DP_E_FIXED can force a fixed value.
  • $L$: average LLM calls per user input. DP_L=0 auto-estimates it from agent_request_count / current_turn_count; otherwise it uses the configured value.
  • $R = E \times L$: expected remaining LLM calls.
  • $\text{avg}$: average input tokens per LLM request, estimated from cumulative stats, defaulting to 4000.
  • $r_t = \max(r^{c+1}, 0.37)$: cumulative retention after this compaction, with a floor to avoid runaway penalties.
  • $M$: max context tokens from MAX_CONTEXT_TOKENS, default 200000.
  • $Q$: DP_QUALITY_PENALTY, default 0.2.

4. What Each Term Means

4.1 ① Future Savings

$$ \text{①} = (R - 1) \cdot P_{\text{cache}} \cdot H / 10^6 $$

After compaction, later requests carry $H$ fewer old-history tokens. Those tokens are usually already in the provider prompt cache, so the savings are priced at $P_{\text{cache}}$.

4.2 ② Cache Miss

$$ \text{②} = (S + K) \cdot (P_{\text{input}} - P_{\text{cache}}) / 10^6 $$

After compaction, the prefix becomes “new summary + retained recent messages”. That prefix is new on the next request, so the model accounts for the difference between full input price and cached input price.

4.3 ③ Compaction Request Cost

$$ \text{③} = [P_{\text{cache}}(V + H) + P_{\text{input}}L_{\text{instr}} + P_{\text{out}}S] / 10^6 $$

The summary call reuses the normal request prefix. The fixed prefix $V$ and dropped history $H$ are mostly charged at cache-hit price; only the appended summary instruction, about $L_{\text{instr}}=70$ tokens, is full-price input. The summary output $S$ is charged at output price.

4.4 ④ Information Distortion Penalty

$$ \text{④} = \beta \cdot (1 - r_t) \cdot R \cdot \text{avg} \cdot P_{\text{input}} / 10^6 $$

Summaries lose detail. The implementation estimates the loss using cumulative retention after this compaction, $r_t=\max(r^{c+1},0.37)$, and scales it by the expected future token volume $R \times \text{avg}$.

4.5 ⑤ Quality Improvement Savings

$$ \text{⑤} = Q \cdot P_{\text{input}} \cdot \frac{(V + T)^2 - (V + K)^2}{M \cdot 10^6} $$

This is the newer positive term. It models the fact that very long contexts degrade answer quality. Compacting shortens the working context from $V+T$ to $V+K$, reducing the expected cost of retries, misses, and corrections. Because it is an incremental improvement, it is added to the net benefit.

5. Parameters

Parameter Default Description
DP_P_INPUT 3.0 Uncached input price, $/MTok
DP_P_CACHE 0.30 Cache-hit input price, $/MTok
DP_P_OUT 15.0 Output price, $/MTok
DP_V 5000 Fixed prefix tokens: system prompt, tools, old summary, etc.
DP_S 500 Estimated summary output tokens
DP_L 0 Average LLM calls per user input; 0 means auto-estimate
DP_BASELINE_E 8 Baseline expected remaining user-input turns
DP_E_FIXED 0 Fixed E override; values greater than 0 skip dynamic estimation
DP_R 0.8 Single-summary information retention rate
DP_BETA 0.03 Information distortion penalty coefficient
DP_QUALITY_PENALTY 0.2 Long-context quality decay penalty coefficient
DP_MIN_KEEP_RATIO 0.12 Minimum ratio of message lines to retain
MAX_CONTEXT_TOKENS 200000 Max context limit, used for forced compaction and quality-term normalization

6. Implementation Details

  • Per-line tokens are approximated as (byte_length + 3) / 4 + 1.
  • min_keep keeps at least 3 lines and respects DP_MIN_KEEP_RATIO.
  • The model evaluates every k ∈ [min_keep, NR].
  • Automatic compaction only happens when best_benefit > 0.
  • The chosen cut point is moved back to a user-message boundary.
  • Context above 90% of the limit, plan_clear, and plan_confirm can force a fallback compaction.

7. Summary

The bash-agent compaction model is not simply “compact when over threshold”. It weighs:

  • cache-priced savings from dropping old history;
  • one-time cache invalidation from the new summary prefix;
  • the summary request cost itself;
  • information loss from repeated summarization;
  • quality gains from shortening an overlong context.

This avoids compacting repeatedly near the boundary while still compacting more aggressively when the working context itself is likely hurting answer quality.


Related article: Cache-Aligned Summarization
Source code: compact_dp.awk, agent.sh

Clone this wiki locally