One reproducible runner for four long-term-memory benchmarks, plus an
LLM-guided multi-round retrieval method. LoCoMo, LongMemEval, EverMemBench,
and SubtleMemory now share the same staged ADD → SEARCH → ANSWER → JUDGE
workflow, resume model, metrics, and run manifest. The new
llm_multiround search method iteratively recalls episode blocks, asks a
separately configurable decider to retain core evidence and issue follow-up
queries, and returns a bounded final context without a cross-encoder.
Added
- LLM-guided multi-round episode retrieval.
POST /api/v2/memory/searchacceptsmethod = "llm_multiround"for user
memory. Each round independently fuses BM25 and vector candidates for the
current sub-queries with RRF, then uses the decider to select core evidence
and identify remaining gaps. Existing search response fields are unchanged;
decider failures are recorded in structured logs and optional trace dumps. - Independent decider configuration. The new
[decider]section selects
the model, endpoint, timeout, request extras, retry policy, and loop tuning.
Empty connection fields inherit[llm], preserving single-model setups. - Unified benchmark harness. One runner and four adapters cover LoCoMo,
LongMemEval, EverMemBench, and SubtleMemory with explicit dataset configs,
resumable stage artifacts, deterministic run identities, shared IR metrics,
decider preflight checks, and trace validation that rejects degraded
multi-round runs before reporting scores. - Cascade snapshot control.
POST /api/v1/cascade/quiesceand
POST /api/v2/cascade/quiescedrain the Markdown-to-index queue and stop the
cascade subsystem until restart. Startup switches can disable all cascade
work or only the filesystem watcher for managed read-only workloads.
Changed
- OME attempts are bounded. One strategy attempt now has a 1,800-second
default wall-clock timeout and follows the existing retry/dead-letter path on
timeout. Environment settings can tune concurrency or disable the timeout. - SQLite pool saturation fails visibly. Pool size, overflow, checkout
timeout, recycle, and pre-ping are explicit settings; exhausted pools now
raise after a bounded wait and emit saturation diagnostics instead of waiting
indefinitely. - Extraction LLM transport is configurable.
[llm]now exposes request
timeout and provider-specific SDK arguments while retaining the previous
60-second default.
Fixed
- Permanent embedding request errors no longer retry. HTTP 400, 401, 403,
404, 413, 414, and 422 responses are classified as rejected input or
configuration and leave the retry loop; timeouts, 429 responses, and server
errors remain retryable. - EverMemBench owner mapping no longer changes profile extraction. The
adapter uses one stable synthetic owner per topic for both ingestion and
search while preserving each real participant's name in message metadata.
The production Profile cluster/direct extraction paths are unchanged.
Upgrade
pip install --upgrade everos # or: uv sync- No storage migration or index rebuild is required. Existing search methods
keep their prior behavior unless callers explicitly select
llm_multiround. - Deployments with a healthy OME strategy that legitimately runs for more than
1,800 seconds must raiseEVEROS_OME_RUN_TIMEOUT_SECONDSor set it to0/
offbefore upgrading.
Full changelog: v1.3.0...v1.3.1