Skip to content

EN Course 06 Compaction and Long Sessions

lloydzhou edited this page Jun 1, 2026 · 2 revisions

Compaction and Long Sessions

Long sessions need compaction, but compacting too early can waste tokens and break cache reuse. bash-agent treats compaction as an economic decision instead of a simple threshold.

agent_compact_context() {
    local trigger=${1:-auto} total_lines keep_lines drop tmp_dropped dropped_messages summary_response
    keep_lines=$(store_conv_dp_decision \
        "$(store_stats_get current_turn_count)" "$(store_stats_get agent_request_count)" \
        "$(store_stats_get compact_request_count)" "$(store_stats_get total_input_tokens)") || true
    [[ -n "$keep_lines" ]] || keep_lines=0
    total_lines=$(store_conv_line_count)
    # drop old lines, summarize them, keep recent turns
}

Cache-Aligned Summary

Summary requests reuse the same system prompt and tool prefix shape. That keeps provider cache hits high even when older conversation lines are compacted.

The summary output is written to summary.txt, then included in later system prompts as stable context.

Dynamic Decision

The DP model balances:

  • future cache savings
  • cache miss cost
  • summary generation cost
  • information loss
  • long-context quality penalty

For the full formulas, see:

Practical Trigger Examples

The decision is data-driven. A few simplified examples:

Situation Observed signal Expected action
Short fresh session low total input tokens, low turn count do not compact
Long session with repeated cache misses high tokens, miss cost growing compact old turns and keep recent lines
Long session but high cache reuse high tokens, stable cache hits delay compaction

store_conv_dp_decision computes a keep-window (keep_lines) from current stats instead of using a fixed threshold.

Prompt Shape Before/After Compaction

Compaction drops old conversation lines, but keeps the outer prompt shape stable:

[system sections]
[instruction-files]
[summary]
[recent conversation window]
[current user input]

The key invariant is structural consistency. Only the summary body and recent window content change.

Failure and Fallback

If summary generation fails, runtime safety still comes first:

  • keep running the current turn without injecting a broken summary
  • preserve recent conversation lines as the minimum working context
  • retry compaction later when conditions are better

This fallback avoids corrupting context while still allowing long sessions to proceed.

Next

Tool Calling explains how model requests become runtime actions once the session and prompt layers are in place.

Clone this wiki locally