Skip to content

Remembra 5.7.1 — retrieval, SDK and cold-start fixes

Latest

Choose a tag to compare

@lacrous lacrous released this 30 Sep 16:11
· 29 commits to main since this release

Three fixes, all found by installing 5.7.0 from the registry and using it, plus two
cold-start races found by building a reproduction rather than re-running a suite.

Full compatibility statement: docs/v5.7.0-compatibility.md.
Findings and measurements: docs/v5.7.0-audit.md.

Fixed

  • Roadmap §37's budget is reachable through the typed SDK. SearchInput — which
    the SDK's SearchOptions is an alias of — never gained budget, so a typed caller
    could not set a bound. Adding it exposed a second half: the SDK serialises query
    parameters with String(value), so a nested object became the literal string
    "[object Object]" and the server ignored it. The SDK accepted the argument, the
    type allowed it, and the bound was never applied. Nested parameters are now
    flattened as parent.child. The response side had the same gap in reverse: the
    server has returned budget since §37 shipped and SearchResponse did not declare
    it.

  • Non-ASCII queries now reach the lexical path. extractQuery filtered query
    terms with /^[a-z0-9]+$/u, so every non-ASCII query produced zero terms, no
    keyword list was built, and retrieval fell back to the vector path alone — nothing at
    all without an embedding provider. A CJK, Russian, Greek, Hangul or Thai query
    returned an empty result set with no error and no warning, which is
    indistinguishable from a query with no match. Both halves of retrieval now share one
    tokeniser, so they cannot drift apart again.

    One ASCII behaviour changes, pinned by a test: a query with internal punctuation now
    contributes its alphanumeric parts, so don't contributes don and matches a
    document containing don't.

  • latest N <query> now parses. TEMPORAL_RE anchored its alternation with $,
    so each branch had to consume the entire query — latest 3 errors fell through and
    was tokenised as the ordinary words latest and errors. Both branches now carry a
    trailing remainder, with the digit count still required so latest news … stays an
    ordinary query.

    This restores parsing, not meaning. The qualifier's values are read in exactly
    one place, where they switch the recency multiplier from 0.5 to 2. latestCount
    still does not limit to N, and before/after still filter nothing. Measured, with
    four documents of decreasing age and all four matching:

    search(pool, "incident review")   -> d1, d2, d3, d4
    search(pool, "latest 2")          -> d1, d2, d3, d4
    search(pool, "before 2026-08-01") -> d1, d2, d3, d4
    
  • Two cold-start races in the batch idempotency ledger. Found by building a
    reproduction, because the symptom was a single MATRIX-02 failure that passed on every
    rerun.

    • A peer's integrity-key staging file made the next start fail. link() needs its
      source to exist, so publishing the key necessarily has a window in which a
      .claims.key.<pid>.<hex>.tmp is visible in the claim directory, and the
      constructor's allowlist rejected that name. Only a staging file this code creates
      is now tolerated, and the constructor waits for its publisher. Anything else is
      still refused immediately.
    • The ledger identity was published with open(O_EXCL) followed by a write, so
      the file existed empty before being filled; a peer reading in between refused to
      start with "ledger identity is invalid". Same defect as the zero-byte integrity key
      fixed in 5.6.0, which survived because that fix was applied to claims.key and not
      to claims.identity. Now published with link() too.

    Measured over 60 rounds of 16 concurrent cold starts: "unrelated files" goes from
    routine (hundreds of failures) to zero; "ledger identity is invalid" from ~1 in 300 to
    zero.

    Two rarer symptoms in the same family remain and are not fixed — they are in code
    this release does not touch, and a per-symptom wait is what produced five
    manifestations in the first place: "ledger database exists without its identity" and
    "database could not be opened" (SQLite contention), about 3 occurrences across
    ~250,000 cold starts. This store has now produced six manifestations; the real fix is
    a single atomic initialisation, and this release does not pretend six patches reach it.

Known limitations

  • The temporal qualifier parses but does almost nothing beyond the recency multiplier.
    See audit S8.
  • The benchmark set contains no temporal query and no non-ASCII document, so the gate
    cannot see two of the fixes above. The unit tests cover them. Both are gaps in the
    gate, recorded rather than papered over.
  • Two cold-start symptoms listed above remain in the ledger.

Gates

940 test suite · 142 security matrix · 52 recovery matrix · Python SDK ·
47 documentation files · 23 benchmark scenarios · 0 audit findings · 0 npm
vulnerabilities. Full release:check passed for @hilbras/remembra@5.7.1 on
Node 18.20.8.