Skip to content

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 04 Sep 00:55
· 169 commits to master since this release

Fixed

  • mem recall --hint-format silently discarded a positional query, returning identical output for any two searches -- recall is declared .command("recall [query]") (src/cli.ts), so the query is a positional, not a --flag. buildHintFormat/buildHintFormatUnsafe (src/integration-seam.ts) hardcoded query: "" in their call to retrieve(), so mem recall "some query" --hint-format --root . and mem recall "zzzz nonexistent xyzzy" --hint-format --root . produced byte-identical output on the same store, with nothing indicating the query had no effect. The plain (non-hint-format) recall path already treats this exact silence as a defect it refuses to ship -- it prints note: query matched no fact text -- showing most recent instead precisely so "the reader is shown unrelated facts with no cue that their query contributed nothing to the ordering" -- but that guard sat after the point where --hint-format already returned, so the same silence shipped on the agent-facing path, where it is worse: the consumer surfaces display strings verbatim into an LLM's context with no way to notice they are unrelated to what was asked.

    The positional now flows all the way through: recall's action passes the query into the new HintFormatOptions.query, buildHintFormatUnsafe forwards it to retrieve() in place of the hardcoded "", and retrieve()'s existing BM25 ranking (previously dormant on this path, since a query-less call scores every candidate 0 and falls through to recency) now actually reorders the emitted TGMEM lines by relevance. A query is a ranking input, never a filter -- matching nothing reorders, it does not narrow, consistent with plain recall's documented contract. An absent or empty query is unchanged byte-for-byte: SessionStart's query-less call still gets the original recency-only behaviour, since BM25 ties every candidate at 0 with no query terms. The incompatibleFlags guard this file's previous entry added for the positional (recall "<query>" --hint-format exiting 1) is removed now that the query has a real effect to honor instead of an effect to error on.

  • Every installed memory-capture instruction taught agents to write facts that can never become ground truth -- the "## Memory" prose mem init writes into CLAUDE.md (CLAUDE_CODE_CLAUDE_MD_BODY) and into the shared AGENTS.md block for Codex/Copilot CLI/Copilot VS Code (AGENTS_MD_SHARED_BODY, both in src/wiring.ts) told an agent to run mem remember with --kind/--subject/--value, but never mentioned --anchor at all. Three-valued freshness (affirmed/unverified/contradicted) is this project's most distinctive design idea, and an anchor is the only way a fact ever earns affirmed ground truth instead of surfacing caveated forever -- yet an agent following the installed instructions literally, which is the whole point of installing them, had no way to know the capability existed. Verified against a real store on this machine: 3 facts, 0 anchored, every recalled line reading (unverified, YYYY-MM).

    Both constants now add a short --anchor line naming the real predicate set (file-exists, file-absent, file-newer-than, glob-exists, git-tracked, newest-of), one concrete example, and the constraint that the anchor path must stay inside --root (no .., no absolute path, which v0.3.2 already rejects at capture with exit 1). docs/integrations/{claude-code,codex,copilot-cli,copilot-vscode}.md and this repository's own CLAUDE.md are updated to match byte-for-byte, verified by the existing structural doc-consistency test that compares each doc's fenced block against what install() actually writes to disk.

  • A fully-consumed retrieval budget was reported as headroom, making the seam's one deterministic degradation handle depend on the clock -- buildHintFormat (src/integration-seam.ts) computed truncated = elapsed > budgetMs, so a budget of N milliseconds counted as exceeded only after N had passed, never on reaching it. At the real 150ms budget that distinction is invisible, but retrievalBudgetMs: 0 is the handle callers and tests use to force the budget-exhausted contract on purpose, and under a strict > a zero budget reported a healthy response whenever the work finished inside a single millisecond. A ten-fact, anchor-free retrieval on a fast runner does exactly that: elapsed reads 0, 0 > 0 is false, and the caller that allotted no time at all was handed a full hint set as though the budget had been honored.

    This surfaced as an intermittent Linux-only CI failure of tests/unit/integration-seam.test.ts's own budget test -- the one whose comment says forcing the budget to 0 exercises the degradation path "so the degradation contract is pinned rather than inferred from a flake". It failed on ubuntu-latest for commit 9fd6f2e and passed on the next commit with the test code byte-identical, which is what identified it as a boundary condition rather than a regression: the assertion was a coin flip decided by how fast the runner completed the retrieval.

    truncated is now elapsed >= budgetMs: the budget is the time available, so having consumed all of it is already an overrun, and a zero budget is unconditionally exhausted regardless of platform or load. Nothing in production passes 0 (options.retrievalBudgetMs ?? RETRIEVAL_BUDGET_MS only substitutes for nullish, and the suite's non-truncating constant is an hour), so the change is confined to the boundary itself. The new regression test freezes Date.now so elapsed is pinned to exactly 0 on every platform, making this the boundary case rather than a race that happens to land on it -- it fails against the old > on Windows, where the original flake never reproduced.

Added

  • Recall's lexical ranking (computeBm25Scores / tokenize, src/retrieval.ts) did not stem terms, so a plural, gerund, or otherwise inflected query term shared no token with a differently-inflected document term -- tokenize was a bare lowercase split on non-alphanumerics, so "commits" and "commit", or "test" and "testing", were entirely different terms to BM25 even though a human reader would call them the same word. This mattered more once --hint-format started honoring a real query (the entry above): a caller typing the natural form of a word it wants ("running the tests") could silently miss a fact phrased in a different inflection ("test runner") for no reason a caller could predict or work around, other than memorizing the store's exact wording.

    tokenize now runs every purely-alphabetic token through an inline Porter stemmer (M.F. Porter, "An algorithm for suffix stripping", 1980) before returning it, applied identically at index time (computeBm25Scores building each document's term frequencies) and query time (the same function scoring the query against them) -- there is no persisted index to go stale, since BM25 recomputes term statistics fresh over the candidate pool on every retrieve() call, so "stemmed index, unstemmed query" (or the reverse) cannot happen by construction. A token containing a digit (e.g. "es6") is left untouched, since the algorithm's vowel/consonant rules are meaningless applied to digits and mangling it would only lose information. Written inline rather than taken as a dependency: this project deliberately removed zod and sqlite-vec for being unreachable weight, and the algorithm is small and fully specified by Porter's paper. Verified against porter-stemmer (jedp/porter-stemmer, MIT), a long-established independent implementation, across the paper's own worked examples plus this project's own vocabulary (commits/commit, testing/test, running/run) -- every pair matched.

  • mem recall --hook-stdin: the recall is driven by what the user just asked, read from the hook's own stdin envelope -- Claude Code hands every hook a JSON object on stdin (session_id on every event; the submitted text in prompt on UserPromptSubmit), not env vars, and the previous SessionStart-only wiring had no way to see a prompt at all. --hook-stdin (--hint-format only) reads and parses that envelope inside mem itself -- the installed hook stays a one-liner with no jq or other dependency on the user's PATH -- takes the envelope's session_id as the session id and probes prompt, then user_prompt, then message for the first non-empty string to use as the recall query (src/hook-envelope.ts). It fails open on every axis: a TTY, an unreadable or slow-to-close pipe, non-JSON, a non-object, or fields of the wrong type all degrade to "no query / no session" and the recall proceeds unranked with exit 0, because a hook that exits non-zero or prints a parse error into an agent's context is worse than one that returns unranked facts. An explicit --session-id <id> overrides the envelope's; the envelope's prompt outranks a positional query.

  • recall_log table and --delta: repeated recalls in one session no longer re-send facts the agent already has -- every --hint-format recall that knows its session id (from --hook-stdin or --session-id) now records the ids it actually emitted in a new recall_log(fact_id, session_id, surfaced_at) table (src/storage.ts; CREATE TABLE IF NOT EXISTS, so a store created by 0.3.2 or earlier is migrated on first open with nothing else touched). The write is best-effort: a failure to log is reported on stderr and never fails the recall. Nothing is logged under --stable, which exists to make output deterministic for tests. mem epoch --gc prunes rows older than 30 days alongside its existing audit_log rotation, reported as a new pruned_recall_log_rows= field on its summary line.

    --delta (--hint-format only) then emits only facts not already logged for this session id -- plus any logged fact that matches the current query (non-zero retrieval score), because a session's host compacts context over time and a fact surfaced as filler on an early prompt must still arrive when it becomes the exact answer later; zero-vs-nonzero is used rather than a threshold because BM25 scores are unbounded and corpus-dependent, and the kind boost multiplies so it cannot lift a zero -- and the header becomes TGMEM/2 delta=1. Delta is opt-in per call, and that is the load-bearing constraint: TGMEM/2 is a closed grammar whose consumers drop off-grammar lines, and budget exhaustion already returns empty, not partial, precisely because a partial block is byte-indistinguishable from a complete one -- so a consumer that never asked for a delta never receives one, and the one that did can see from the header that it got one. --delta without a session id (no --hook-stdin, no --session-id) is a usage error, not a silent full response; under --hook-stdin an envelope that arrives without a session_id degrades to the full set with a stderr note, since a hook must not exit non-zero over that. The filter runs before the per-kind caps, so a session drains the next-best unseen facts across prompts rather than getting an empty block as soon as the top-ranked ones have all been sent once. HintFormatResult gains a delta boolean; the grammar comment in src/integration-seam.ts documents the optional header flag.

  • mem init claude-code now installs a UserPromptSubmit hook, and both hooks read the stdin envelope -- the installer writes mem recall --hint-format --hook-stdin --delta --root "$CLAUDE_PROJECT_DIR" under hooks.UserPromptSubmit, behind the same command -v mem ... || true guard as the existing SessionStart hook, so it still exits 0 with no stderr on a machine without mem. The SessionStart command gains --hook-stdin too: it carries no prompt, so that recall stays query-less and recency-ordered exactly as before, but its session_id is now the same one the later prompt-time deltas subtract from -- without it the first prompt would have re-sent everything the session opener already surfaced. Both hooks are stamped and reference-counted like before (src/wiring.ts), so re-running init upgrades an existing stamped SessionStart command in place, a pre-existing hand-written hook with the same command still aborts with a conflict error, and mem uninstall claude-code removes exactly the two stamped entries and the containers it created, restoring a pre-existing settings.json byte-for-byte. docs/integrations/claude-code.md shows the new two-event block and the structural doc-mirror test now checks both events.

  • tokenize (src/retrieval.ts) now drops standard English function words before stemming, at both index and query time -- with --delta re-sending any already-surfaced fact that scores non-zero for the current prompt, a prompt like "what is the plan for today" re-sent most of a store on the/is/for/a alone (measured on a 6-fact store: 4 of 7 realistic unrelated prompts re-sent 2/6 facts; a more function-word-heavy store leaked 5/6), which was the difference between delta working and not. The list (STOPWORDS, exported) is a small conventional set of articles, prepositions, auxiliaries, pronouns, conjunctions, and interrogatives; no domain terms, and deliberately not a long published list. Negations (no, not, nor, never, none, neither, nothing, without, cannot, dont, cant, wont) are kept as content words on purpose, because this store holds negative preferences and corrections, and n't contractions are expanded to ... not before the apostrophe split so "don't"/"can't"/"won't" keep their negation. A prompt made entirely of stopwords leaves no query terms and falls through to the existing no-query behaviour (recency order, fully suppressible under --delta); a fact whose text is mostly stopwords is still indexed by its content words. Nothing is persisted from tokenize (scores are computed per call), so existing stores are unaffected. Plain mem recall ranking changes too: on the npm run eval fixture set (BM25 in isolation, not end-to-end) query-stem moved from p@8 74.7% / nDCG@8 95.6% to 77.3% / 98.7%, deterministic across runs. mem recall --hint-format additionally honours a TOKEN_GOAT_MEM_RETRIEVAL_BUDGET_MS test/advanced override for the 150 ms soft budget, so subprocess-level tests are not hostage to runner load.