Repository navigation
Deep research is the worst case for served context ceilings, and compaction is where the question quietly changes #5896
Replies: 1 comment 1 reply
|
Thanks for raising this. A fluent report that answers a subtly different question after compaction or a stage handoff is a failure mode worth testing in DeerFlow. The current summarization middleware preserves the latest real user message by ID, and A useful next step would be a redacted reproduction with the configured provider/model endpoint, original question and constraints, stage handoffs, compaction points with estimated token counts, and the specific place the final report drifted. That could support a regression asserting that the writer receives the original question and required constraints. I would validate effective context limits and truncation behavior per endpoint/API rather than assume one fixed ceiling across providers. I'd avoid logging the raw question or source text by default; hashes or explicitly redacted traces would be safer for diagnostics. |
Uh oh!
There was an error while loading. Please reload this page.
Disclosure: I build Grunz, a hosted chat and coding agent on open weights. Different product, but deep-research flows and long agent runs fail in the same place, and this one is badly under-documented.
Deep research is the worst case for the context ceiling, and the ceiling is not the one on the model card
The window advertised on a model card is a property of the weights. The window you are actually served is whatever your provider configured for that endpoint. Most hosted endpoints serve around 32K regardless of the advertised figure. A few reach 256K. I have not found one genuinely serving the 1M numbers that get quoted.
It fails silently — no error, no truncation notice, the front of the context is simply evicted and the model continues sounding confident.
A deep-research flow is the most exposed thing you can build on top of that, because the whole design accumulates: search results, fetched pages, per-source notes, an evolving outline. The very thing that makes the output good is what buries the original question.
The concrete failure: by the time the writer stage runs, the research plan has been through several rounds of summarisation, and the specific question the user asked has been compressed into a topic. The report that comes back is well-sourced, well-written, and answers a slightly different question than the one asked. Nothing errors. Nobody can point at the step where it went wrong, because the step where it went wrong was a compaction nobody logged.
What actually helped
Pin the user's question verbatim through every compaction and every stage handoff. Not summarised, not "captured in the plan" — re-emitted in full. Then have each stage restate it before producing output; when the restatement drifts, stop. This is small and it prevents the entire class of failure above.
Log compaction as a first-class event — what was dropped, what was kept, tokens before and after, which stage. Without it, debugging a bad report means re-reading the whole trace and guessing.
Budget in remaining context, not in steps. "Search up to N times" carries no information about whether there is room left to actually write the report. "You are at 80% of the window and have not written the draft anywhere durable" does.
Treat every stage handoff as a lossy compaction. In a multi-agent graph, each edge is another chance to drop the constraint, and it is crossing a boundary where nothing is checking that the goal survived. More agents does not fix the context problem; it multiplies the number of places it can occur.
Two smaller things
Per-provider, not per-model. Because different nodes can point at different models and providers, the effective ceiling varies inside one run. A planner on a genuinely long endpoint feeding a writer on a silently-32K one gives you a coherent plan and a report that drops half of it, with nothing in either log explaining why. Worth probing and surfacing the observed ceiling per configured endpoint rather than trusting a declared value — sending a deliberately oversized request and reading the error text usually leaks the real number.
For anyone running open weights here: ablation costs instruction-following and output-format adherence before it costs knowledge. Prose stays good while structured output drifts off-spec, which in a graph where nodes parse each other's output is exactly the failure that bites. Perplexity will not warn you.
Happy to go deeper on any of it.
All reactions