Compaction cannot rescue an oversized session — the summarization request itself exceeds the provider request-size limit (HTTP 413) #7626
Replies: 3 comments 2 replies
|
Your read of The mechanism, confirmed
The step that decides the outcome: 413 is classified by body wording, not by status
So a 413 lands on — is the fallback string, used only when the response body carried no Why that is the whole story:
This is the actionable part for maintainers: a 413 with no body message is undecidable between "the model's context is too small" and "the request exceeded a transport limit" — and the default branch (400/413 → Your question, answered
To bound it, no. No request-size guard exists (above), and the config surface is token-only:
One real nuance: a failed compaction still shrinks the session durably. The model-free prune phase lands before the summarization call, and the overflow path says so explicitly ( There is one way out today, and it is not a fork at the tail: branch from an early And why your workaround works, since it is worth knowing rather than just applying: the trigger is a fraction of the window you declare (
That reframes your suggestion 2: it is not only "derive the trigger from the transport limit too", it is that the declared window is currently trusted as the request-size authority, and a second authority (bytes) is missing. Your suggestion 1 (chunked/hierarchical summarization) is the real fix and belongs in What a plugin can and cannot do here — and one honest boundaryI scanned this against the public seams rather than assume:
Your suggestion 4 (refuse or warn when forking a corpus already past the limit) is the cheapest of the four and is also plugin-mountable — |
|
Thanks — the classification step is the piece my report was missing, and it changes the conclusion. I've corrected the retry framing: I ran your citations against
One divergence: I've already applied two workarounds — branching from an early turn, and lowering the provider model's declared |
|
Thanks — I've installed and mounted 0.2.0 with |
Uh oh!
There was an error while loading. Please reload this page.
Summary. A long session with heavy tool output can reach a state where every request fails with HTTP 413 (payload too large), and automatic compaction cannot recover it: the summarization call replays the compacted region in one request and hits the same limit. The session becomes unusable — no turn succeeds, compaction fails with the same 413, and a fork inherits the corpus (
isSeeded: true), so it fails from its first turn.Environment
0.1.7-alpha.2launched from source, Windows 11 x64, bundled Node 24.17.0deepseek-official, modeldeepseek-v4-flash, declaredcontextWindow: 1000000in the provider configtool/result); decompressed corpus ≈ 56 MB, stored.jsonl.zstd≈ 8 MBObserved
compaction/pruneshadows a few old results (~19K tokens), thencompaction/start→compaction/endwitherror: "DeepSeek Messages request failed (413)", and the routed request ends with the same 413 (turn/end, reasonerror).Suspected cause (from source)
packages/compaction/compaction-basic/src/summarizer.ts:summarizeWithLlm()builds ONEctx.llm.stream()call carrying the replayed conversation prefix (input.messages: system head plus the whole shadowed region) plus the compaction instruction. The replay is intentional (prefix-cache reuse), but there is no chunking and no size guard.packages/compaction/compaction-basic/src/config.ts: budgeting is token-only —thresholdRatiodefault 0.8,retainRatiodefault 0.16,headroomTokensdefault 65,536,maxTokensdefaulting to the headroom.packages/compaction/compaction-basic/src/region.tspicks the region purely by token counts; neither module accounts for request bytes.Impact. Once a session crosses the transport limit there is no way out: automatic compaction cannot run,
/compactpresumably fails the same way, forks inherit the problem, and the only option is to abandon the session — losing continuity for exactly the long-running work the feature exists for.Suggested directions
Question. Is there a supported way today to bound a session's request size, or to recover an oversized session (trim / partial compaction)? If yes, please document it; if not, would a byte-aware trigger plus chunked summarization be acceptable? I can share a redacted event excerpt (event names, sizes, error strings) if that helps.
Workaround for anyone hitting this: lower the provider model's declared
contextWindow(for example to 131,072–262,144) so compaction triggers much earlier and new sessions stay inside the provider's request limit. The oversized session itself stays unusable.All reactions