v0.2.2
Added
- Memory benchmark harness (
eval/benchmarks/) running LOCOMO,
LongMemEval, and BEAM against sage-wiki as the system under test,
using the datasets, prompts, and judging procedure of
mem0ai/memory-benchmarks
(Apache-2.0, vendored with attribution). Each conversation compiles into
its own sage-wiki project, then retrieval runs throughsage-wiki search.
Results with gpt-5 as answerer/judge on scoped samples: LOCOMO 92.0%
@ top-50 (150 q), LongMemEval-S 93.3% @ top-50 (30 q), BEAM 100K
0.691 mean nugget (60 q) — seeeval/benchmarks/REPORT.mdfor the
comparability caveats, which are substantial. - Rate-limit resilience for long runs: a process-wide gate shared by the
LLM client and the search subprocess. One worker's 429 pauses every worker
(exponential backoff,Retry-Afterhonored), rate-limited search degrades
get more retries than permanent ones, and sustained limiting aborts the run
cleanly with resume instructions rather than burning the remaining queue
into failures.
Fixed
eval.pycould not read any real wiki. It hardcoded_wiki/while the
scaffold emitsoutput: wiki, so it exited 1 on every project the current
binary produces; it now resolves the output directory fromconfig.yaml.eval.pyfact-extraction always scored 0% on real wikis — it counted
only bullet lines under## Key claimswhile the summarize pass writes
prose, so a single run reported 0% extraction and 100% "Structural — Key
claims". Counting is now format-agnostic, and the section regex no longer
swallows an empty section into the following heading.- Manifest lock treated Windows contention as fatal. A contended
exclusive-create returnsERROR_ACCESS_DENIEDon Windows (pending-delete or
sharing violation), whichos.IsExistdoes not match — so a routine lock
race aborted the caller'sMutateand silently dropped its update.
Concurrent manifest writers on Windows could lose data.
Changed
- README benchmark numbers now come from real compiled wikis, not from
eval_test.py's synthetic fixture generator. The April 2026 figures
(85.9–86.7%) described the generator's parameters; measured across 10 real
wikis the overall score is 87.4% median with 100% fact extraction.
Updated in all seven READMEs.
Fixed
- CLI
searchnow honors the configured hybrid weights (default
0.7 BM25 / 0.3 vector) and thesearch.ann.enabledsetting — it
previously fused with 1.0/1.0 and always brute-forced. The dead
--scopeflag (parsed, never read) is removed. - Web
/api/searchnow embeds the query and passes the configured
hybrid weights — it previously passed a nil vector, silently running
BM25-only regardless of embedding configuration. - Chunks found only by vector search are hydrated with their real
heading and content before reranking/output — they previously flowed
through fusion as empty passages. - Rerank blending is now safe under partial LLM coverage: candidates
the LLM never scored keep their normalized [0,1] relevance instead of
being coerced to zero, blending operates in normalized space on both
sides (never raw RRF ~0.016 vs LLM [0,1]), and when the LLM scores
fewer thansearch.rerank_min_coverage(default 0.5) of the head the
blend is skipped entirely, keeping RRF order.
Added
sage-wiki reindexrebuilds the chunk index from the documents on
disk using the current chunking config — compiled articles
(concepts/,summaries/,outputs/) and chunk-indexed raw sources
alike, replaced per document (delete-then-insert). No LLM article
writing happens. Re-chunking changes chunk IDs, so old chunk vectors
cannot be kept: without an embedding provider the command stops rather
than empty the chunk-vector leg, and--drop-chunk-vectorsrebuilds the
text index anyway (chunk-level vector search stays off until the next
compile --re-embed).- Soft tag boost:
sage-wiki search --boost-tags a,band MCP
wiki_search{boost_tags}rank documents carrying those tags +3% each
(capped at 15%) without excluding anything — the complement to
--tags/tags, which filter. The boost was specified and implemented
but had no caller until now. search.chunk_overlap_tokens(default 0, recommended opt-in
80, max half ofchunk_size): each chunk after the first repeats the
tail of its predecessor, so a fact straddling a chunk boundary is
retrievable from either side. The default 0 is byte-identical to previous
chunking — upgrading never re-chunks an existing index. Changing the
value takes effect only viasage-wiki reindex; edit the config and
reindex as one step, or the index mixes both chunkings (docs:
search-quality.md § Chunk overlap).
Changed
-
Web
/api/searchscoreis now the normalized [0,1] fused score, not
the raw RRF score (~0.016 scale) — a client thresholding on it needs new
thresholds. New field:source_date(unix seconds; omitted when the
document has no known origin date). -
sage-wiki search --config <path>now fails when that file cannot be
loaded instead of silently searching with default weights and no
vectors; auto-discovered config still degrades with a warning.
--expand/--rerankfail when no LLM client can be built, and
sage-wiki reindexrefuses to run against an unloadable config. -
Query-term stopwording is corpus-adaptive: above 100 documents, terms
matching more than 20% of documents are dropped from the lexical query,
and both the document and chunk legs prune the same term set (they now
probe the same corpus, which also removed the chunk leg's per-term
COUNT(DISTINCT)join — the single largest cost in unified search). -
Every search surface now runs the unified retrieval pipeline
(MCPwiki_search, CLIsearch, web/api/search, and the TUI —
sage-wiki querymoved in M2): chunk-level and document-level hits fuse
with the configured weights, the ontology graph contributes as a third
channel, and dated documents get the recency tie-breaker. Result
ordering changes on all of them. Results gainFinalScore,
GraphRank,SourceDateandAliasOf(web:source_date) alongside
the existing fields, and MCP/CLI gain per-callchannels,expandand
rerankoptions — the LLM stages stay OFF unless asked for. The TUI ran
BM25-only with default weights before; it now uses the configured
weights, vectors and graph like every other surface. -
search.pipeline: legacypins any surface back to the previous
doc-level path if the new ranking is disruptive; the value is validated,
so a typo is rejected rather than silently resolved tounified. -
Result hydration is a single batched query (
EntryStore.GetMany)
instead of one lookup per result document — it was the unified
pipeline's dominant per-query cost, and removing it puts the unified
path at parity with (slightly under) the legacy doc-level path's
latency on a 1k-entry corpus despite searching both chunks and
documents. -
Search entry points now apply the
trust.include_outputsrule
(MCPwiki_search, CLIsearch, web/api/search, TUI, and
hub search):output:
documents — LLM-generated answers auto-filed back into the wiki — are
excluded unless the mode admits them (truealways,verifiedonly
once confirmed). The default isfalse, so by default these surfaces
no longer return outputs. Previously only the Q&A path enforced this,
and an agent searching the wiki could read what a Q&A answer would
refuse to cite. Settrust.include_outputs: trueto restore the old
search behavior. -
sage-wiki queryretrieval was rewritten as a unified weighted
fusion (20260728-search-upgrade M2): document- and chunk-level hits
now both contribute (agreement across granularities ranks higher),
the configured hybrid weights apply on this path for the first time,
multi-query expansion variants sum instead of taking the best rank,
and rankings will shift accordingly. -
The ontology graph now joins retrieval ranking as a third fused
channel (sage-wiki querypath): query terms seed entities (alias
links included), a depth-2 traversal with per-relation weights
(contradicts1.1,cites0.7) ranks their neighborhoods, and the
results fuse atsearch.hybrid_weight_graph(default 0.2). Articles
reachable only through the graph can now surface; results carry
their graph rank and analias_ofnote when reached via an alias.
An empty ontology costs nothing (byte-identical results). -
All search surfaces (query, MCP, CLI, web, TUI, hub) inherit two
lexical upgrades through the shared query builder: on corpora over
100 documents, query terms matching more than 20% of documents are
pruned (corpus-adaptive stopwording), and entry matching now weights
the id/article-path columns 3× over body content (title-proxy boost;
Postgres schema migration v6 rebuilds the search vector). Result
rankings on every surface change accordingly. -
README diagrams refreshed (architecture, compiler pipeline,
interfaces) — higher-resolution replacements for the three PNGs. -
Translated READMEs regenerated to full parity with the restructured
English README (zh, ja, ko, vi, fr, ru): same 22 sections, localized
internal anchors, translated bash-fence comments with byte-identical
commands, verbatim config identifiers and numbers. The temporary
restructuring banners are removed; the standing may-lag marker stays.