Repository navigation
Replies: 2 comments
|
Added detail: the distillation runner should be customizable to use high-context, cost-effective models (like Gemini / Antigravity harness or custom ACP endpoints) so users don't burn expensive frontier model tokens just summarizing history. |
0 replies
This comment was marked as spam.
This comment was marked as spam.
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Problem
When working on large agentic tasks (especially multi-step refactors or research), sessions often hit model rate limits or session usage limits (e.g., Anthropic 5-hour quota).
By the time the limit window resets:
Currently, the user must manually open a new thread and summarize the exact files, worktrees, and next steps to avoid burning tokens.
Proposed Solution
Add an "Auto-Fork to Fresh Thread" action (or banner when a provider reports
quota_exceeded/session_limit):"Continue in New Thread with Compact Context". Clicking it opens a fresh thread pre-populated with the distillation, preserving the clean state while paying ~300 tokens instead of 150k.All reactions