Skip to content

v4.3.0

Choose a tag to compare

@Korrnals Korrnals released this 15 Sep 22:15
· 578 commits to main since this release

Changed

  • Search v2 query semantics — FTS v2 builder, project soft fallback, graph leg, embedding_id stamp (issue #313, ADR-0029) (src/mnemos/storage/sqlite_store.py, src/mnemos/manager.py, src/mnemos/models.py, src/mnemos/storage/vector_store.py, src/mnemos/cli/main.py, tests/{test_search_v2_golden,test_search_v2_scope_fallback,test_search_v2_graph_leg,test_embedding_id_backfill}.py — new, tests/test_security.py) — the M15.2 whole-input quoted phrase (the injection hardening) made multi-token queries adjacency phrases ("GWS конвейер" → 0 live hits vs 4 for the per-token AND) and disabled prefix matching (конвейер 49 rows vs конвейер* 77 — inflected RU/EN forms invisible). The v2 builder keeps the hardening and removes the side effects: every token is individually quoted with a builder-owned prefix star ("tok"*), AND-joined (cap 8 terms), with a de-hyphenated OR-alternative per hyphenated token (("release-trigger"* OR "release"*) — unicode61 splits on hyphens, so the bare identifier can never match the split index); AND-empty multi-token queries retry ONCE with the OR join (bm25-ranked, logged). Scope drift (release-pipeline vs releases-pipeline) no longer silently zeroes scoped searches: the A9 pre-RRF predicate is unchanged, a NEW OUTER retry without the scope surfaces the rows TAGGED (SearchResult.project_scope_fallback, backward-compatible field) and audited (search_stats()["project_scope_fallback_total"], new counter distinct from cross-project; explicit status= drill-downs are NOT retried — the same status policy holds on the retry, no junk resurfacing). New graph leg v1: after RRF fusion the top-limit fused ids expand 1 hop along memory_edges (supersedes, both directions — get_incoming_edges added), edge rows not already fused are appended with the deterministic decay (1-alpha)/(rrf_k + 2*anchor_rank) (same FTS weight at a strictly deeper position — never outranks its anchor; the naive 1/(rrf_k+2*rank) DOES at alpha 0.5 and was caught by the golden decay test), carry via_graph=True, pass the SAME status/quarantine (ADR-0019 §5 — an edge is never a side door)/refined-only (§4) gates, capped at limit extra rows. memories.embedding_id (NULL for 1644/1644 live rows — a diagnostic trap; the vector leg resolves by id and worked) is now stamped by upsert_embedding on every write and closable for existing rows via mnemos backfill-embedding-ids (idempotent id-join against vectors.db, dry-run default, --apply to write; rows whose vector is gone stay NULL — the column never lies); the production-DB backfill run is deliberately NOT part of this slice. Snowball stemming is deliberately deferred (FTS rebuild migration) — prefix terms are the zero-migration morphology fix; recorded as the follow-up in ADR-0029. Tests: golden suite re-derives every live-probe defect as a semantic assertion (multi-token far-apart AND, inflected prefix, both hyphen spellings, drift-tagged fallback with counter, superseded sibling via graph, injection battery on the builder output, single-token no-regression proven as a prefix superset of the old exact phrase), plus the M15.2 escaping tests updated to the v2 shape and a builder-level injection class (output always executes; no operator syntax; stars are builder-owned).

  • search.hybrid_alpha default 0.7 → 0.5 — balanced RRF fusion (quick win, issue #300) (src/mnemos/config.py, config.example.yaml + config.container.yaml, docs EN+RU (incl. cli-reference env defaults), e3 run manifests) — one constant: the RRF fusion weight now balances the FTS/vector legs. At alpha 0.7 the vector leg structurally subordinated any FTS-only match (an FTS-rank-1 hit scored 0.3/61 < a pure-vector rank-1 at 0.7/61); at alpha 0.5 they tie, so FTS-rank-1 matches stop drowning by construction — no special-case code. Measured on both embedder regimes (probe, issue #300): governance top-5 68→73/96 nano (+5.2pp, the full B1 recoverable set) / 68→69/96 lexical (+1.04pp); knowledge recall@5 +3.9pp nano / +3.7pp lexical; G-neg top-5 composition byte-identical (zero displacement). Per-call hybrid_alpha= overrides (manager/SDK/MCP) unchanged. Event-driven S1/S1m re-baseline per ADR-0020 (composition-algorithm change): recall@5 0.8637→0.902967 (S1 hybrid) / 0.863002→0.902269 (S1m nano); model fingerprint unchanged — fusion, not embedder. e3 run manifests now pin retrieval.hybrid_alpha in the content-addressed core (probe finding 6) so a future default re-tune is visible to e3 content-addressing; recorded runner-1 runs stay valid as history (version-gated verify).

  • Search v2 short-token guard — degenerate tokens no longer collapse bm25 idf (issue #314, closes via #317; ADR-0029) (src/mnemos/storage/sqlite_store.py, tests/test_search_v2_golden.py, tests/test_security.py, benchmarks/baselines/{s1.json,BASELINE.md}, docs/project/adr/0029-search-v2-query-semantics.md) — 1-char tokens and RU/EN stopwords (_FTS_STOPWORDS frozenset, exact lowercase entries) matched nearly every row, collapsing bm25 idf to 0 and degrading any AND query containing one to LIMIT rows ordered by id-tiebreak noise. The guard runs between tokenisation and _FTS_TERM_CAP: 1-char tokens are dropped; stopwords are dropped unless the token is fully uppercase (acronym exemption — IT, QA, DB, CI, ML, GWS, and AND-as-literal survive; Cyrillic isupper() acronyms included); identifier-shaped tokens (v2, x1, p0) survive structurally (>=2 chars, never alpha stopwords — no digit-scanning code); an all-degenerate query keeps its original tokens (never-empty fallback — the builder must not introduce silent zeros); dropped tokens consume no term-cap budget. Injection safety untouched: tokens still flow through the same per-token quoting, no new user text reaches MATCH raw. S1/S1m baselines re-recorded per ADR-0020 (semantics change is permanent, the measured shift is real): reference recall@5 0.9366 → 0.9409, recall@10 0.9503 → 0.9642 (planted appearances 202 → 221); S1m recall@5 0.9275 → 0.9484; corpus/model fingerprints unchanged. Tests: golden suite section 8 (13 cases — builder pins, acronym/digit survival, never-empty whole-list fallback, cap interplay, manager-level recall) plus the test_security quoted-only pin updated to the guarded token list.