Repository navigation
Replies: 2 comments
|
Two things from metering long tool loops on a different harness, both about group A. Round masking with a sliding boundary defeats prefix caching. With keep-last-N, every new round rewrites the result from N rounds back, and that message sits in front of the rounds you kept. Provider caches match on an exact prefix, so each round bills the kept rounds again as fresh input. Masking in steps (wait until K rounds are eligible, then replace them in one pass) keeps the prefix stable between steps. Truncation at insertion is worth more than masking later. A tool result is sent again on every round after it arrives, so its cost is size times rounds remaining. In one of my sessions a 104k token read at step 3 was read back as 19.6M tokens, while a read of the same size at step 182 cost 133k. That argues for option 2 ahead of option 1. On group B, state that stores pointers to retained originals has degraded more gracefully for me than state that stores copies. I maintain a reference architecture with the cost model written up at https://github.com/jimy-r/agent-workspace-architecture/blob/main/PATTERNS.md#18-position-is-price--a-token-costs-more-the-earlier-you-add-it Does option 1 mask at a sliding boundary today? Drafted with Claude Code and reviewed by me before posting. |
|
Thanks for putting this proposal together. I agree that context compaction is a foundational and important capability for Agent. I’m currently focused on the 0.4 release, but I’ll take a closer look and join the discussion when I have some bandwidth. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Hi all. I'd like to propose adding context compaction to Flink Agents and agree on its shape here before opening PRs. Below: the failure, what we found when we measured the obvious fixes, the strategies we propose, and our questions.
The failure
A ReAct-style agent running as a Flink Agents job (PyFlink 2.3.0,
ActionExecutionOperator, the built-in chat model action,OpenAIChatModelSetup, gpt-4o-mini). One input event, one tool returning 40 rows (~2.6k tokens). To reach the limit quickly the setup forces a tool call on every round (tool_choice="required", sequential calls), so this loop cannot end on its own; the un-forced agent in the second bullet below hits the same wall. Every round the built-in loop re-sends every earlier tool result. After 48 rounds and 53 seconds:ChatResponseEvent; retrying an unchanged oversized request cannot fix it, and depending on the durable-execution configuration recovery either replays the recorded failure or repeats the work (the run above disables restarts). Prompt caching does not help, cached tokens still count.Why there is no bound today
MemoryObject.setcan replace stored history (there is noremove, see the TODO inchat_model_action.py), but the framework provides no conversation-compaction policy.The existing controls (event-logger payload truncation, the routing judge's budget, and since #1040 detection of truncated output) do not impose a total input-context budget on the ordinary chat loop.
What the cheap fixes preserve
We measured the cheap fixes before proposing anything, on two synthetic workloads with known answers and on the data-driven agent above; numbers on request. Full history gives wrong answers while it still fits, then fails hard. In these experiments, masking old tool results reduced prompt growth but could remove information needed later. Summarisation achieved higher answer recall than masking, with additional model calls and changes in task behaviour. Inside Flink, masking, per-result truncation and a round cap took the data-driven agent to 41 rounds with no provider error, but it never produced an accepted report: avoiding overflow does not guarantee task completion.
State instead of a transcript
We also ran the 70-event key from the failure above with a small working state per key instead of a transcript, updated by the model on each event and sent in place of history. The request stays at one event plus one state object, about 2.4k tokens, for all 70 events, against a 400 on event 61 for the transcript. The risk moves to state quality: a free-form state kept the BR whitelist but dropped its ticket id, which cannot be recovered from the state alone; prompting with an explicit schema kept both complete fact and ticket pairs; with both facts correctly stored, gpt-4o-mini passed two of four probes (one answer was incorrect, another omitted the required ticket) and gpt-4o passed all four in this single run. These are illustrative runs; the state arms used additional instructions and a larger output budget than the transcript arm. A working state gives no verbatim recall, so tasks that need exact historical details need retained originals and a retrieval path.
Proposal: nine strategies, four groups, all off by default
Masking and truncation change only the model-visible messages and preserve the active request's stored tool-call payloads and tool_call_id pairing; the event log keeps what its truncation setting allows. The round cap changes control flow, group B changes stored state, and native compaction adds continuation state.
A — deterministic trimming, no model call (a prototype exists in Java and Python with tests and docs; it needs a hardening pass before a PR)
chat.tool-results.keep-last-rounds: keep the last N rounds' results verbatim, replace older ones with a one-line placeholder (tool, arguments, omitted size). Assistant tool-call messages and placeholders stay, so the prompt still grows slowly (about 53 tokens per round in our run); this trims growth rather than bounding it. Peers: LangChainClearToolUsesEdit, Anthropicclear_tool_uses, MicrosoftToolResultCompactionStrategy.chat.tool-results.max-chars: trim one oversized result in the middle. Precedent: core Flinkcontext-overflow-action.chat.max-tool-rounds: fail through [api][plan][java][python] Represent terminal chat failures in response events #1149's path after N rounds without a final answer. Peers: OpenAI Agents SDKmax_turns, Confluentmax_iterations.B — state instead of transcript, no separate compaction call
MemoryObject.remove(setalready overwrites), an append primitive, and a documented pattern with the example above (typed object per key, schema validation, hard size cap). Stored in Flink managed keyed state and included in configured checkpoint recovery (not exercised in the run above). Venue: [Tech Debt][Umbrella] Review and refine Flink Agents APIs for 0.4 #1055, [Tech Debt][API][Chat Model] Review the chat model invocation flow in custom Actions #1086, [Tech Debt][API][ChatMessage] Review ChatMessage responsibilities and data model #1056.C — a model compresses
D — a model decides (direction only, built on the routing judge)
For 5–9: durably recorded summariser and judge outcomes can be reused on replay; calls without a recoverable recorded outcome may execute again. The current action API has no background-compaction lifecycle, so adding one needs explicit coordination with per-key state and recovery. Not proposed: sliding window as primary, retrieval as compaction, token pruning or KV-cache methods; code-mode tools and sub-agents are patterns to document, not options.
Why not just …
max-tool-roundsis its loop-level sibling.Questions
chat.tool-results.keep-last-roundsthe right name and level (execution option vs chat-model setup field)?max-tool-roundsoffer afinishmode (one last call without tools) besides failing?remove+appendonMemoryObjectwith a documented pattern enough for group B in 0.4, or should the framework offer a typed working-state helper?References
EventsCompactionConfig; Microsoft Agent Framework compaction; Confluent Streaming Agentstokens_management_strategy(TRIM / SUMMARIZE pastmax_tokens_threshold, off by default) andmax_iterations; Snowflake Cortex Agentsagent:compact(preview); core Flinkmax-context-size/context-overflow-action.All reactions