You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Compactor trait (re-exported from rig_core::memory) + CompactingMemory<M, P, C> adapter and a TemplateCompactor
reference implementation. Where DemotingPolicyMemory only observes messages a policy truncates out of active history, a Compactorsubstitutes them: it derives a single Message-shaped Artifact from the evicted prefix (and, optionally, the previous
summary), and CompactingMemory splices that artifact at the front
of the loaded history. The resulting prompt shape is [summary_message, ...kept_window], with the summary rolling
forward on every load that produces newly-evicted messages — the
canonical recursive-summary pattern for long-running agents.
Concurrent loads on the same conversation_id are serialised at
the compaction seam via an in-flight gate: only one load at a time
invokes the compactor; others observe the gate and immediately
return the previously-stored summary spliced in front of kept,
without re-running the compactor. Watermarks and the carry-over
summary are in-process only — Compactor implementations with
durable side effects (LLM calls, vector-store writes) must
deduplicate. clear drops the carry-over so a freshly-populated
backend re-compacts from scratch. TemplateCompactor is a
zero-dependency, no-LLM rollup useful as a default and for tests;
it produces a TextSummary that converts into a Message::System
with header + previous-summary + per-line role: text body. The
rollup represents out-of-band context about the prior conversation
rather than a turn from any participant, so the system role is the
semantically correct framing across providers. TemplateCompactor
exposes with_max_bytes so long-running conversations can cap the
rolled-up text; when exceeded, the oldest portion of the body is
dropped at a UTF-8 boundary and replaced with a "[…truncated…]"
marker, preserving the most recent context. Note that the spliced
summary sits outside the wrapped MemoryPolicy's budget, so
pairing CompactingMemory with a token-budgeted policy requires a
bounded compactor (or accepting that the loaded prompt may exceed
the policy budget by the artifact size).
DemotionHook trait + DemotingPolicyMemory<M, P, H> adapter and a NoopDemotionHook no-op default. The trait itself lives in rig_core::memory (re-exported here) so any memory backend can
implement it without taking a rig-memory dependency; the composing
adapter lives in this crate. DemotingPolicyMemory calls the hook
with messages that the policy truncated out of active history,
turning eviction into demotion. It tracks per-conversation demotion
watermarks so repeated load calls do not replay the same demoted
messages into append-only long-tail stores. Concurrent loads on the
same conversation_id are serialised at the demotion seam via an
in-flight gate: only one load at a time delivers to the hook;
others observe the gate and return the truncated history without
re-firing. Watermarks are in-process only — DemotionHook
implementations must be idempotent on (conversation_id, messages)
to survive process restarts. Bridges SlidingWindowMemory / TokenWindowMemory to long-tail stores such as MemvidPersistHook
without coupling either crate to the other.
DemotingPolicyMemory::forget(conversation_id) and tracked_conversations() for explicit watermark-map cleanup and
leak diagnostics. Both are infallible: a poisoned internal lock
is treated as "nothing to forget / zero tracked" rather than a
caller-visible error.
MemoryPolicy::apply_with_demoted companion method that reports (kept, demoted). apply remains the required method; the default apply_with_demoted returns (apply(...)?, Vec::new()) so existing
policies keep compiling unchanged. SlidingWindowMemory and TokenWindowMemory override it to populate the demoted prefix that DemotingPolicyMemory hands to the hook.
HeuristicTokenCounter — provider-agnostic, zero-dependency TokenCounter implementation that approximates token cost from
UTF-8 byte lengths (str::len, O(1)). Ships default / openai / anthropic / gemini presets so TokenWindowMemory::new(budget, HeuristicTokenCounter::default())
works out of the box without a tokenizer dependency. Also handles the Message::System variant and tool-call argument payloads. The
configurable ratio is named bytes_per_token to match the
implementation; for ASCII text bytes and characters coincide, and
for non-ASCII text the counter slightly over-estimates, which is the
safe direction for a hard budget.
PolicyMemory<M, P> adapter — wrap any ConversationMemory with a MemoryPolicy and propagate policy failures to the caller as MemoryError::Policy. Hard-fail counterpart to InMemoryConversationMemory::with_filter + IntoFilter::into_filter,
which degrade to identity on policy error.