You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Keep a Claude coordinator's prompt cache warm while it waits for background agents
#14055
When a Claude thread ends its turn while background agents are still running, the main agent makes no requests until a task notification arrives. Claude Code caches the main thread with the 1-hour TTL, and only a read refreshes it. If one background agent runs longer than an hour and nothing else finishes, the next wake-up re-reads the whole context at full price.
In one of my sessions (Claude Code 2.1.283, Opus 5.5, context about 900k tokens), the coordinator slept 42 and 44 minutes between notifications and still hit the cache, with about 16 minutes to spare. A miss there would rewrite about 900k tokens: about $7.20 at API prices, versus about $0.18 to read them.
T3 already has the signal it needs. ProviderSessionReaper skips sessions whose turn has settled while background work runs (thread.backgroundLiveness).
Proposal: an opt-in setting for Claude threads. When the turn has settled, background work is still running, and about 50 minutes have passed since the last API response, T3 sends a short message, for example: "Status check from T3: reply in one line with current progress; don't start new work." The reply refreshes the cache for another hour. It doesn't change the coordinator's flow, since a user message during background work is already a normal path.
The cost is one short turn, and the exchange stays in the transcript. A cheaper fix belongs in Claude Code itself (a cache-only request, proposed in anthropics/claude-code#97758); this would work with the current Claude Code.
Related: #7590 (show each thread's remaining prompt-cache window).
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
When a Claude thread ends its turn while background agents are still running, the main agent makes no requests until a task notification arrives. Claude Code caches the main thread with the 1-hour TTL, and only a read refreshes it. If one background agent runs longer than an hour and nothing else finishes, the next wake-up re-reads the whole context at full price.
In one of my sessions (Claude Code 2.1.283, Opus 5.5, context about 900k tokens), the coordinator slept 42 and 44 minutes between notifications and still hit the cache, with about 16 minutes to spare. A miss there would rewrite about 900k tokens: about $7.20 at API prices, versus about $0.18 to read them.
T3 already has the signal it needs.
ProviderSessionReaperskips sessions whose turn has settled while background work runs (thread.backgroundLiveness).Proposal: an opt-in setting for Claude threads. When the turn has settled, background work is still running, and about 50 minutes have passed since the last API response, T3 sends a short message, for example: "Status check from T3: reply in one line with current progress; don't start new work." The reply refreshes the cache for another hour. It doesn't change the coordinator's flow, since a user message during background work is already a normal path.
The cost is one short turn, and the exchange stays in the transcript. A cheaper fix belongs in Claude Code itself (a cache-only request, proposed in anthropics/claude-code#97758); this would work with the current Claude Code.
Related: #7590 (show each thread's remaining prompt-cache window).
All reactions