Skip to content

Releases: DFKHelper/token-goat-mem

v0.5.2

Choose a tag to compare

@github-actions github-actions released this 04 Oct 17:47

Fixed

  • Project identity. Credentials, query strings and inline config comments are stripped from a remote URL before it becomes an identity, so a token never reaches the store. Equivalent spellings of one remote (scp, ssh://, https://, trailing / or .git, case) reduce to the same identity. A relative local remote (../upstream.git, .\x) gives no identity, since it names a different repository from every clone that resolves it. A drive-letter remote (C:\repos\x.git) is read as a local path rather than an scp host named C. A path fact under a filesystem root (C:\, /) binds correctly.
  • Import. An imported scopeRepo keeps its identity intact (port, subpath and a .git-suffixed subpath included) and has any credential stripped, so it still matches the live project. Imported scope roots are normalized. A captured_at in the future is rejected, captured_at and last_surfaced_at go through a strict ISO-8601 parser that accepts a bare date and refuses malformed or extended-year values, duplicate ids are flagged in the --dry-run plan, and superseded_by edges whose id is not 36 characters are kept.
  • Review and undo. Undoing a rejection restores the fact's previous status, and undoing a rejection or an edit skips usage and refused-secret events rather than treating them as the change to undo.
  • Capture. mem remember on text that matches a pending suggestion promotes it with source_type user. A repeated mem suggest counts as a sighting without writing a source row that only echoes the fact. Sighting excerpts keep the fact's own sentence. mem scan-session no longer scans the output that !cmd bash-mode turns wrap in the user turn.
  • Contradiction and consolidate. A fact that agrees with the winner is never superseded, a pinned project fact is never superseded as a cross-scope duplicate, and mem consolidate compares normalized values, so two spellings of one value are not a conflict.
  • Anchors. valid-until dates are parsed strictly, git-tracked answers for the repository root, and a non-UTF-8 file under file-contains, a reftable HEAD placeholder, and unsupported glob syntax read as unverified instead of contradicted.
  • Storage on Windows. A --root that differs from the stored scope root only in case reaffirms the fact instead of missing it.
  • Backup and restore. mem restore recovers from a live store that is not a readable database by moving it aside to mem.db.unreadable-<timestamp> and naming that file in its output, and retries the swap when a concurrent write leaves the epoch unchanged. The stale-temp sweep only touches mem's own snapshot files.
  • Wiring. A project mem init run from the home directory keeps the user-level hooks, every duplicated shared AGENTS.md block is stripped and collapsed, shared Copilot keybindings survive until the last uninstall, and mem uninstall --all --user skips tools that have no user-level config. The postinstall script also treats npm_config_location=global as a global install.
  • CLI. Malformed numeric options are rejected rather than coerced, and daysAgoIso clamps to the valid Date range. The hint footer leaves out a session id that is not safe to pass back on a command line. mem reflect matches session suggestions by the resolved transcript path. The package entry points match the shipped bundle.

v0.5.1

Choose a tag to compare

@github-actions github-actions released this 04 Oct 04:17

Changed

  • Node.js 22.12 or later is now required. Node 18 and Node 20 are both past end of life, and the current releases of mem's runtime dependencies (better-sqlite3, commander) no longer support them. engines is now >=22.12.0 and the bundle is built for node22. CI runs the full suite, including the tests that drive the built bundle, on Node 22 and Node 24 on both Linux and Windows. This replaces the old Node 18 check, which could only smoke-test the bundle because the test toolchain no longer ran there.
  • better-sqlite3 13 (SQLite 3.53). It ships its native binaries inside the npm package, so npm i -g token-goat-mem no longer downloads a prebuilt binary at install time or, when none matches the running Node, compiles one with node-gyp.

Fixed

  • Unquoted multi-word text is refused instead of silently truncated. mem remember the build runs on node 22 stored the fact "the" and exited 0, and mem recall build runs searched for "build" alone: the argument parser (commander 12) dropped excess arguments without a word. With commander 15 both exit non-zero with too many arguments, and nothing is stored.
  • mem doctor reports hooks that run twice. Claude Code merges project and user hooks and skips a duplicate only when its command text is identical. Since 0.5.0 a Windows user hook launches node on the bundle while a project hook still calls mem, so a project with its own mem hooks ran every event twice and received recall twice, while hook-divergence reported the two levels as matching. It is now ok only when the commands are identical, and warns on a launcher-only difference, naming the events and mem init claude-code as the fix (see below); mem uninstall claude-code would also have stripped the CLAUDE.md block.
  • A project mem init claude-code no longer adds hooks that user-level hooks already cover. The user hooks run in every project, so once ~/.claude/settings.json holds mem's hooks a project install writes only the CLAUDE.md block and removes mem's hooks from the project .claude/settings.json, deleting the file when nothing else is left in it. Without user hooks the project install writes its hooks as before. mem doctor counts a project settings file with no mem hooks as current in that case, and its hook-divergence remedies now name mem init claude-code rather than an uninstall.

v0.5.0

Choose a tag to compare

@github-actions github-actions released this 02 Oct 04:24

Added

  • On Windows, mem init claude-code --user writes hooks that launch node on the installed bundle directly. npm's mem shim is a shell wrapper that adds a noticeable delay to every hook event, so the user-level hook command is now if [ -f "<bundle>" ]; then node "<bundle>" <subcommand> ... || <fallback>; elif command -v mem >/dev/null 2>&1; then mem <subcommand> ... || <fallback>; fi, falling back to the plain mem guard when the bundle is gone. The bundle path comes from the running install, uses forward slashes, and is refused (plain shape kept) if it contains ", $ or a backtick. Project hooks and non-Windows installs keep their exact commands. Re-running init upgrades an old-shape hook in place without duplicating it, uninstall removes either shape, and the PATH pre-flight probes the bundle itself when it is the launcher. mem doctor compares project and user hooks by subcommand and flags, so a launcher-only difference is no longer reported as divergence, and a new launch-time finding warns when mem --version takes over 300 ms, measured only when some installed hook would launch the PATH mem (on Windows it suggests re-running mem init claude-code --user).

  • mem review --promote <id> --subject <key> --value <value> keys a pending fact as it activates it. Keying, activation, and the contradiction pass the new key makes the fact eligible for run in one transaction, and the audit row records keyed <subject>=<value>, so a suggested fact no longer needs a follow-up mem edit to become supersedable. Both flags are required together, only with --promote, and are validated and secret-screened like mem remember's; a contested fact is already keyed and is refused with a pointer to mem edit.

  • mem epoch --json emits the current write epoch as JSON. Outputs { epoch: <number> }. The plain output remains unchanged byte-for-byte for cache invalidation.

  • mem facets --json emits structured facets data for each mode. --list-entities --json emits an array of { term, facts } objects; --fact <id> --json emits { fact, entities, topics }; backfill/--all mode emits { facts, entities, topics } summary. Plain output is unchanged.

  • mem backup --list --json emits snapshot records as JSON. Outputs an array of snapshots, each with path, name, takenAt (ISO 8601), epoch, reason, and size fields. The plain output remains unchanged.

  • mem doctor checks wiring drift for every supported tool, and project-vs-user hook divergence. For each tool in mem init's list whose config exists under --root (new on doctor, default the current directory) or the home directory, a wiring finding is ok when the block matches exactly what mem init <tool> would write now, and warn (remedy mem init <tool>, with --user at user level) when it is outdated, hand-edited inside its markers, or a Claude Code hook lacks an expected event. Tools mem was never wired into are one informational ok line, and doctor still writes nothing. A new hook-divergence finding warns when project and user Claude Code hooks run different mem invocations (both run, so recall is duplicated or inconsistent) and names which to remove or re-init.

  • mem forget <id> --by <winner-id> records which fact replaced the forgotten one. Plain forget retires a fact with no named successor, so mem show and mem export could only say superseded_by: unknown. --by takes the winner by id or unique prefix (same resolver and errors as the id argument), refuses the fact itself or a winner that is not active/pinned, and writes the same Superseded by fact <id> audit detail the contradiction and review paths use, in the same transaction as the forget. mem review's pending may contradict line now ends with a paste-ready mem forget <rival> --by <id>.

  • Retrieval now indexes a decision's rationale for lexical search. computeBm25Scores now appends doc.why to the tokenized document text, so queries matching only a decision's rationale (e.g., "staging memory") can find the decision that records it. This lets "why did we..." questions in a future session surface decisions before relitigating them, without waiting for embeddings to refresh.

  • mem doctor --json and --strict. Every doctor line is now a structured finding
    ({ check, status: ok|warn|fail, message, remedy? }) drawn from a fixed list of check names, and
    the plain report is those messages printed unchanged. --json prints { findings, epoch }.
    --strict exits 1 when any finding is fail (an installed hook whose mem binary is missing or
    too old to run it); without it doctor still exits 0, so a warning such as a store with facts and no
    backup never breaks a script that did not ask.

  • npm i -g token-goat-mem installs the Claude Code hooks. A new postinstall script runs the freshly installed mem init claude-code --user on every global install and upgrade, so recall reaches every project without a manual step and an upgrade refreshes hooks an older mem wrote. It keeps mem init's pre-flight check that the mem on PATH can run those hooks, never fails the install (anything short of success prints mem init claude-code --user as the fix), does nothing on a local install, and is skipped with TOKEN_GOAT_MEM_SKIP_HOOKS=1. It also leaves the hooks alone when run as someone other than the home directory's owner (root under sudo npm i -g, which would otherwise leave a root-owned ~/.claude/settings.json), and stops an init that hangs after two minutes (TOKEN_GOAT_MEM_POSTINSTALL_TIMEOUT_MS). npm is the only package manager that runs it: pnpm, yarn, bun and --ignore-scripts installs need mem init claude-code --user by hand.

  • The store is backed up outside the mem home. Snapshots land in ~/.mem-backups (override: TOKEN_GOAT_MEM_BACKUP_DIR), so deleting ~/.mem -- or ~/.claude, which never held the store -- no longer loses every fact. Opening an existing store takes an automatic VACUUM INTO snapshot at most once a day and only when it has changed (the newest 14 are kept), and always before a pending schema migration (once per store state, so a migration that keeps failing does not copy the store on every open). Only the open that wins a claim file in the backup directory takes the daily snapshot, so hooks opening the store from several processes at once copy it once; a claim or partial copy left by a process that died is cleared after an hour. Each snapshot is named for the epoch of the copy itself, never a read taken before it. mem backup takes one by hand and mem backup --list lists them; mem restore <snapshot> checks and migrates a private copy of the snapshot first -- it must pass integrity_check, hold mem's own facts columns, and not come from a newer mem -- then saves the store it replaces as a pre-restore snapshot and swaps the rows in one transaction on the live connection, advancing the epoch past both stores. A write that lands between that snapshot and the swap is caught and the swap retried, so the pre-restore snapshot always holds exactly what was replaced, and the restore never prunes the automatic snapshot it is restoring from. A snapshot that cannot be written never fails the command that opened the store; mem doctor gains a backups: line that reports it instead.

  • mem scan-session can file a correction. Its trigger table yielded fact, preference and
    decision and had no shape that produced a correction, while the block mem installs tells the
    agent to persist exactly those three kinds plus corrections. The scanner was structurally unable to
    find one. Two openers now do: an explicit correction: prefix, matching the existing decision:
    and rule: entries, and the reversal forms of that's wrong / that's outdated / that's no longer true. Bare no, and actually, are deliberately not among them. They open any negative
    answer -- "no, that test is fine" reverses nothing -- and the pending queue is sorted by sighting
    count, so a false positive restated across sessions would climb it. Only shapes that assert the
    reversal in their own words qualify; anything subtler is what mem remember --kind correction is
    for.

  • A cross-process concurrency test. mem has no daemon: every command is a short-lived process
    contending for one WAL database under BEGIN IMMEDIATE, and no in-process test can exercise that,
    because a single Database handle serialises everything by construction.
    tests/bundle/concurrency.test.ts runs six workers of the built bundle against one store that does
    not exist yet, so schema creation races too. Each worker interleaves unique remembers, a
    remember of one shared sentence, and recalls. The test asserts no write is lost, the shared
    sentence stays one row with every later statement audited as a reaffirm, and nothing reports
    SQLITE_BUSY. Verified by opening the database with a zero busy timeout: recall's surfaced-marking
    write then reports a locked database, and the test fails.

  • mem log: the store-wide audit timeline. mem show <id> could read one fact's audit trail,
    but only once you knew which fact to ask about; "what did the agent change in my memory this
    week?" had no answer. mem log lists every audit row newest first, each prefixed with a pasteable
    short fact id, with --fact, --event (exact name or capture-style family), --age-days,
    --limit, and --json. --fact also resolves ids gc has hard-deleted, because audit rows
    are kept for 180 days and superseded facts for 90, and it reports a prefix shared by a live and a
    deleted fact as ambiguous instead of showing only the live one. mem show's history block now
    renders through the same formatter.

  • mem reflect: resolve pending suggestions while the agent still knows what it meant.
    mem scan-session files every durable-sounding sentence as pending a...

Read more

v0.4.1

Choose a tag to compare

@github-actions github-actions released this 16 Sep 14:22

Fixed

  • A store that could not be opened was reported as a project with no memory. buildHintFormat
    wrapped the whole of buildHintFormatUnsafe in one catch, and openStorage is called inside it, so
    a permissions error, a WAL lock or a schema mismatch produced bytes identical to an empty store: the
    bare TGMEM/2 header and nothing else. The hook commands mem installs end in || true, so nothing
    downstream surfaced it either, and the failure mode of the one channel this tool exists to keep
    reliable was indistinguishable from its quietest success. The open now has its own catch and its own
    error type, so the three cases -- unreadable store, nothing to recall, some other internal fault --
    are told apart rather than collapsed. An unreadable store emits a footer clause saying so and still
    exits 0, because failing open is the right behaviour and silence about it was not. The clause points
    at mem doctor, which opens the same store and will fail the same way: that is deliberate, and the
    wording says only that doctor shows the underlying error, which is exactly what it does. Promising
    a command that fixes this would repeat the refusal loop this project fixed one release ago.
  • mem scan-session dated every fact from the moment of the scan. The scanner read a transcript
    entry's text and never its timestamp, and the capture call omitted capturedAt entirely -- so a
    statement made in March and scanned in September was stored as having been said in September. Under
    the Stop and PreCompact hooks the drift is seconds and harmless, but --transcript <path> is a
    first-class flag, and a rescan of an archived session mis-dated everything it produced by the whole
    age of the transcript. That is not cosmetic: preference confidence decays from captured_at, the
    --stale cutoff is measured from it, recall breaks recency ties on it, and the review queue is
    ordered by it. mem suggest and mem import --from-md already distinguished when a thing was said
    from when mem stored it; the scanner was the one capture path that did not. It now reads the entry's
    timestamp through the same validator, which already refuses a future date. Anything absent,
    malformed or hostile falls back to the scan time without failing the run -- a transcript mem cannot
    date is the situation that existed until now, not an error. The source row's stored_at still
    records the scan: said-at and stored-at are genuinely different columns.
  • A source excerpt could omit the sentence it was evidence for. Excerpts were truncated from the
    head at 600 characters, and for mem scan-session the raw material is the whole user turn --
    deliberately larger than the extracted sentence. A turn longer than the cap whose durable statement
    came last therefore produced a sources row that did not contain the fact at all, and mem review
    printed it under that fact as the reviewer's grounds for promoting or rejecting it. Evidence that
    does not contain the claim is worse than no evidence: it invites a decision on the strength of an
    unrelated fragment. The truncation window is now centred on the fact's own sentence, located through
    the same normalizeFactText the store uses everywhere else so collapsed whitespace cannot defeat
    the search, with a marker on whichever side was cut. A raw excerpt that does not contain the fact
    falls back to head truncation, and the screening order is untouched: the untruncated text is
    screened first, and a screened-positive excerpt still yields no source row and still captures the
    fact.
  • A query that matched nothing was answered with recent facts and no mention of it. When a
    non-empty query produces no lexical hit and no other rank list has anything to say, every fact ties
    at zero and the caps fill with whatever is newest. The seam's own comment calls these filler. Under
    the installed hooks this is mostly masked, since --delta suppresses what was already sent, but a
    plain mem recall with a query and --hint-format is a documented command, and it handed a host a
    page of facts under a query none of them matched. The footer now says so. The condition is
    deliberately not "no result matched": matchedQuery is computed from BM25 alone by design, so a
    query carrying real embedding signal reads false on every row, and claiming nothing matched there
    would put a false statement into the wire output. Retrieval now reports whether any rank list ranked
    anything at all, and the clause is gated on that, on a non-empty query, and on there being a fact to
    qualify.
  • An exported supersession edge did not survive being imported. mem export recorded that a fact
    was superseded but not what superseded it, so after a round trip mem show printed superseded_by: unknown -- a caveat about information the store had held and thrown away, given that fact ids
    survive import intact. Export now carries the edge for superseded facts, and import re-establishes
    it. The winner is resolved once the whole file has been inserted rather than row by row, so a file
    naming a successor that appears further down its own list still links; a per-row resolve would have
    silently dropped exactly that case. A named winner that resolves to nothing in the store is ignored
    without writing a dangling edge and without failing the import, and the existing caveat then stands
    -- now meaning a genuinely untraceable supersession rather than a limitation of the format. The
    field is optional in both directions, so the export schema version is unchanged and an older reader
    is unaffected.
  • Nothing mem ever printed told anyone how to say a fact helped. mem used shipped, recall_log
    carried a used_at column for it, and no output on any path named the command -- so in a real store
    that column is null on every row, and every consumer of the signal was reading evidence that could
    not be produced. mem consolidate --stale gated on "never marked used" against a mark nobody could
    make; usefulness counts broke ranking ties that were never actually broken. The recall footer now
    carries a ready-to-run mem used <ids> --session-id <id> clause whenever the response is logged
    under a session and actually emitted a fact -- not when either is missing, since an invocation
    naming a row that was never written earns the exact "was never surfaced in session" refusal this is
    meant to stop producing. Footer text is free prose in the wire grammar precisely so a clause can be
    added here without bumping TGMEM/2: an unknown header version fails open to no hints at all, so
    paying a protocol bump for one footer clause would cost every consumer every fact-line. Separately,
    --stable no longer suppresses recall_log writes. It reorders output for reproducible tests,
    which is no reason for the store to forget what it showed.
  • mem init never upgraded an install that was already there. It wrote its marker, saw a mem hook
    already in settings.json and left the entry as it found it -- so an install predating a hook flag
    or a hook event kept the older command string forever, and re-running mem init after upgrading mem
    did nothing to fix it. On this machine a July install carried only SessionStart, without
    --hook-stdin: no recall_log row had ever been written and mem scan-session had never run once,
    a dead feature that read as a code defect rather than an install-age symptom. mem init now
    recognises an unstamped hook as its own by matching mem's invocation shape -- same guard wrapper,
    same subcommand for that event, any flags in between -- and adopts it, rewriting the command to the
    current one and stamping it so uninstall still reverses exactly what install wrote. A hand-written
    entry that merely mentions mem does not match the shape and is left alone. An adoption is reported
    in the install output rather than performed silently, because absorbing another install's hook is
    the kind of thing a user should hear about from the command that did it.
  • mem consolidate --stale asked a lifetime question of evidence that rotates. The gate read
    "never surfaced by recall", over a recall_log that mem epoch --gc prunes at 30 days. A fact
    surfaced steadily for a year and last read 31 days ago had no surviving row to prove it, so it
    qualified as never-read and was proposed for supersession on the strength of a deletion. The query
    now asks whether the fact was surfaced since the cutoff, which is a question the retained window
    can actually answer, and --help and the summary line say "unsurfaced since then" rather than
    "never surfaced". Usefulness decays with the same rows by the same design, everywhere it is read;
    permanence has its own mechanism, mem pin, which this query excludes by construction.
  • mem import --from-md re-filed facts the store already had. Its dedup key is source location
    plus text, which correctly catches re-importing an unchanged file, and catches nothing about the
    same sentence arriving by another door. A statement captured by mem scan-session, or typed with
    mem remember, would be filed again as a fresh pending candidate the moment it also appeared in a
    markdown note -- so the two capture paths disagreed about what the store knew, and mem review got
    a question already answered. Import now also checks the store-wide text index, scope-gated by the
    same isBoundToRoot rule recall and scan-session use, so an identical fact bound to an unrelated
    project cannot suppress a candidate here. Such a bullet is reported as skipped (already known),
    distinct from skipped (duplicate), because "you imported this file before" and "mem learned this
    elsewhere" are different things to tell someone.
  • mem init --user quietly configured a project instead. For a tool whose only target is a file
    in the repository -- codex, copilot-cli -- --user ...
Read more

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 04 Sep 00:55

Fixed

  • mem recall --hint-format silently discarded a positional query, returning identical output for any two searches -- recall is declared .command("recall [query]") (src/cli.ts), so the query is a positional, not a --flag. buildHintFormat/buildHintFormatUnsafe (src/integration-seam.ts) hardcoded query: "" in their call to retrieve(), so mem recall "some query" --hint-format --root . and mem recall "zzzz nonexistent xyzzy" --hint-format --root . produced byte-identical output on the same store, with nothing indicating the query had no effect. The plain (non-hint-format) recall path already treats this exact silence as a defect it refuses to ship -- it prints note: query matched no fact text -- showing most recent instead precisely so "the reader is shown unrelated facts with no cue that their query contributed nothing to the ordering" -- but that guard sat after the point where --hint-format already returned, so the same silence shipped on the agent-facing path, where it is worse: the consumer surfaces display strings verbatim into an LLM's context with no way to notice they are unrelated to what was asked.

    The positional now flows all the way through: recall's action passes the query into the new HintFormatOptions.query, buildHintFormatUnsafe forwards it to retrieve() in place of the hardcoded "", and retrieve()'s existing BM25 ranking (previously dormant on this path, since a query-less call scores every candidate 0 and falls through to recency) now actually reorders the emitted TGMEM lines by relevance. A query is a ranking input, never a filter -- matching nothing reorders, it does not narrow, consistent with plain recall's documented contract. An absent or empty query is unchanged byte-for-byte: SessionStart's query-less call still gets the original recency-only behaviour, since BM25 ties every candidate at 0 with no query terms. The incompatibleFlags guard this file's previous entry added for the positional (recall "<query>" --hint-format exiting 1) is removed now that the query has a real effect to honor instead of an effect to error on.

  • Every installed memory-capture instruction taught agents to write facts that can never become ground truth -- the "## Memory" prose mem init writes into CLAUDE.md (CLAUDE_CODE_CLAUDE_MD_BODY) and into the shared AGENTS.md block for Codex/Copilot CLI/Copilot VS Code (AGENTS_MD_SHARED_BODY, both in src/wiring.ts) told an agent to run mem remember with --kind/--subject/--value, but never mentioned --anchor at all. Three-valued freshness (affirmed/unverified/contradicted) is this project's most distinctive design idea, and an anchor is the only way a fact ever earns affirmed ground truth instead of surfacing caveated forever -- yet an agent following the installed instructions literally, which is the whole point of installing them, had no way to know the capability existed. Verified against a real store on this machine: 3 facts, 0 anchored, every recalled line reading (unverified, YYYY-MM).

    Both constants now add a short --anchor line naming the real predicate set (file-exists, file-absent, file-newer-than, glob-exists, git-tracked, newest-of), one concrete example, and the constraint that the anchor path must stay inside --root (no .., no absolute path, which v0.3.2 already rejects at capture with exit 1). docs/integrations/{claude-code,codex,copilot-cli,copilot-vscode}.md and this repository's own CLAUDE.md are updated to match byte-for-byte, verified by the existing structural doc-consistency test that compares each doc's fenced block against what install() actually writes to disk.

  • A fully-consumed retrieval budget was reported as headroom, making the seam's one deterministic degradation handle depend on the clock -- buildHintFormat (src/integration-seam.ts) computed truncated = elapsed > budgetMs, so a budget of N milliseconds counted as exceeded only after N had passed, never on reaching it. At the real 150ms budget that distinction is invisible, but retrievalBudgetMs: 0 is the handle callers and tests use to force the budget-exhausted contract on purpose, and under a strict > a zero budget reported a healthy response whenever the work finished inside a single millisecond. A ten-fact, anchor-free retrieval on a fast runner does exactly that: elapsed reads 0, 0 > 0 is false, and the caller that allotted no time at all was handed a full hint set as though the budget had been honored.

    This surfaced as an intermittent Linux-only CI failure of tests/unit/integration-seam.test.ts's own budget test -- the one whose comment says forcing the budget to 0 exercises the degradation path "so the degradation contract is pinned rather than inferred from a flake". It failed on ubuntu-latest for commit 9fd6f2e and passed on the next commit with the test code byte-identical, which is what identified it as a boundary condition rather than a regression: the assertion was a coin flip decided by how fast the runner completed the retrieval.

    truncated is now elapsed >= budgetMs: the budget is the time available, so having consumed all of it is already an overrun, and a zero budget is unconditionally exhausted regardless of platform or load. Nothing in production passes 0 (options.retrievalBudgetMs ?? RETRIEVAL_BUDGET_MS only substitutes for nullish, and the suite's non-truncating constant is an hour), so the change is confined to the boundary itself. The new regression test freezes Date.now so elapsed is pinned to exactly 0 on every platform, making this the boundary case rather than a race that happens to land on it -- it fails against the old > on Windows, where the original flake never reproduced.

Added

  • Recall's lexical ranking (computeBm25Scores / tokenize, src/retrieval.ts) did not stem terms, so a plural, gerund, or otherwise inflected query term shared no token with a differently-inflected document term -- tokenize was a bare lowercase split on non-alphanumerics, so "commits" and "commit", or "test" and "testing", were entirely different terms to BM25 even though a human reader would call them the same word. This mattered more once --hint-format started honoring a real query (the entry above): a caller typing the natural form of a word it wants ("running the tests") could silently miss a fact phrased in a different inflection ("test runner") for no reason a caller could predict or work around, other than memorizing the store's exact wording.

    tokenize now runs every purely-alphabetic token through an inline Porter stemmer (M.F. Porter, "An algorithm for suffix stripping", 1980) before returning it, applied identically at index time (computeBm25Scores building each document's term frequencies) and query time (the same function scoring the query against them) -- there is no persisted index to go stale, since BM25 recomputes term statistics fresh over the candidate pool on every retrieve() call, so "stemmed index, unstemmed query" (or the reverse) cannot happen by construction. A token containing a digit (e.g. "es6") is left untouched, since the algorithm's vowel/consonant rules are meaningless applied to digits and mangling it would only lose information. Written inline rather than taken as a dependency: this project deliberately removed zod and sqlite-vec for being unreachable weight, and the algorithm is small and fully specified by Porter's paper. Verified against porter-stemmer (jedp/porter-stemmer, MIT), a long-established independent implementation, across the paper's own worked examples plus this project's own vocabulary (commits/commit, testing/test, running/run) -- every pair matched.

  • mem recall --hook-stdin: the recall is driven by what the user just asked, read from the hook's own stdin envelope -- Claude Code hands every hook a JSON object on stdin (session_id on every event; the submitted text in prompt on UserPromptSubmit), not env vars, and the previous SessionStart-only wiring had no way to see a prompt at all. --hook-stdin (--hint-format only) reads and parses that envelope inside mem itself -- the installed hook stays a one-liner with no jq or other dependency on the user's PATH -- takes the envelope's session_id as the session id and probes prompt, then user_prompt, then message for the first non-empty string to use as the recall query (src/hook-envelope.ts). It fails open on every axis: a TTY, an unreadable or slow-to-close pipe, non-JSON, a non-object, or fields of the wrong type all degrade to "no query / no session" and the recall proceeds unranked with exit 0, because a hook that exits non-zero or prints a parse error into an agent's context is worse than one that returns unranked facts. An explicit --session-id <id> overrides the envelope's; the envelope's prompt outranks a positional query.

  • recall_log table and --delta: repeated recalls in one session no longer re-send facts the agent already has -- every --hint-format recall that knows its session id (from --hook-stdin or --session-id) now records the ids it actually emitted in a new recall_log(fact_id, session_id, surfaced_at) table (src/storage.ts; CREATE TABLE IF NOT EXISTS, so a store created by 0.3.2 or earlier is migrated on first open with nothing else touched). The write is best-effort: a failure to log is reported on stderr and never fails the recall. Nothing is logged under --stable, which exists to make output deterministic for tests. mem epoch --gc prunes rows older than 30 days alongside its existing audit_log rotation, reported as a new pruned_recall_log_rows= field on its summary line.

    --delta (--hint-format only) then emits only facts not already logged for this session id -- plus any logged fact that matches the current query (non-zero retrieval score), because a session's host compacts context over ...

Read more

v0.3.2

Choose a tag to compare

@github-actions github-actions released this 02 Sep 23:52

Fixed

  • Contradiction detection split one Windows project into two buckets, on the same machine, in the same session -- bucketKey (src/contradiction.ts) keyed a fact's contradiction bucket on subject + scope + the raw scopeRoot string, and scopeRoot is path.resolve(root) at capture time, which preserves whatever drive-letter/segment case the invoking shell happened to report. On Windows, cd /c/Projects/foo and cd /c/projects/foo are the same directory but resolve to differently-cased strings, so two mem remember calls for the same project from two differently-cased shells produced two distinct bucket keys instead of one. Reproduced against the built bundle: recording package-manager=pnpm from C:\Projects\token-goat-mem and package-manager=npm from C:\projects\token-goat-mem left both facts active in mem list -- two directly contradictory facts about the same subject in the same project, neither superseded, both surfacing as ground truth simultaneously. That is the exact failure the contradiction layer exists to prevent, and every existing contradiction test used a consistently-cased root, so nothing exercised the mismatch.

    bucketKey now case-folds its root component through the same win32-only normalizePath that retrieval.ts already applies when comparing paths for scope binding (moved to src/pathUtils.ts so both modules import one definition instead of drifting comment-for-comment; retrieval.ts already imports contradiction.ts, so re-exporting the function from retrieval.ts itself would have created a cycle). The fold stays win32-only, matching anchors.ts's FS_CASE_INSENSITIVE: macOS is case-insensitive by default but supports case-sensitive APFS volumes, so folding unconditionally would trade a missed match for a false one there.

  • A committed .claude/settings.json broke every session for a collaborator without mem on PATH -- mem init claude-code --root . writes a SessionStart hook running mem recall --hint-format --root "$CLAUDE_PROJECT_DIR" into .claude/settings.json, and that file is not gitignored, so it is ordinarily committed and shared. A collaborator who clones the repo without mem installed hit bash: line 1: mem: command not found (exit 127) at the start of every Claude Code session -- the fail-open contract this seam otherwise guarantees (README: a missing or broken mem never blocks a session) had no effect here, because the failure happened one level up, in the hook invocation itself, before mem ever got a chance to fail open.

    The installed command is now command -v mem >/dev/null 2>&1 && mem recall --hint-format --root "$CLAUDE_PROJECT_DIR" || true: the command -v gate skips the call entirely when mem is missing, and the trailing || true forces exit 0 either way, since a bare && guard alone would still exit 1 (and could still be surfaced as a failed hook) when mem is absent. docs/integrations/claude-code.md's hook example and the settings.json this installer writes are covered by a structural test that compares them directly, so the two cannot drift. Deliberately out of scope: switching the install target to settings.local.json would sidestep the sharing problem at the source, but that is a larger design decision left for a separate pass.

  • mem import --from-json silently accepted a scoped fact with no binding -- validateJsonFact checked scope but never cross-checked scopeRoot, so a hand-edited export, an export from a version predating scope_root, or a trimmed migration file could import a project/path-scoped fact with a missing or empty-string scopeRoot, reported as imported. Downstream the two consumers disagreed about an empty string: retrieval.ts's isBoundToRoot excluded it (so plain mem recall never returned it -- an invisible fact), while integration-seam.ts's isInScope only excluded null, so scopeRoot: "" fell through to resolvePath(""), resolving to process.cwd() -- putting the fact in scope for --hint-format from any project whenever --root equaled the cwd, which is the normal case. A project fact could leak into every project's hint block.

    validateJsonFact now rejects a non-global fact with no non-empty-string scopeRoot, naming the index and the problem. A global fact's stray scopeRoot is normalized to null rather than rejected, since global ignores it everywhere it is read. isInScope now excludes an empty/whitespace-only scopeRoot the same way isBoundToRoot already did, closing the leak for facts already in a store today. Absolute paths are still not required: README already documents verbatim cross-machine scopeRoot behavior for --from-json, and requiring isAbsolute would break that contract.

  • --scope path was documented but unreachable -- capture.ts's applyOptionalFields set scopeRoot = resolve(root) for every non-global scope, so mem remember "auth.ts owns migrations" --scope path --root . -- the only invocation the README suggested -- bound the fact to the project directory, not the file. A path fact behaved exactly like a project fact: mem recall --hint-format --root . --context-files src/other.ts returned it for every file in the project, not just the one it was supposedly bound to. remember, suggest, edit, and import --from-md now accept --path <file>, resolved against --root and stored as scopeRoot when --scope path is given. --scope path without --path, or --path without --scope path, now exits 1 with a message naming the missing flag instead of silently binding to the wrong thing.

  • mem list could not tell projects apart -- its summary line showed [kind/status] with no project/path binding, so in project B, mem list --scope project showed project A's decisions indistinguishably, and a user could run mem forget/mem pin against the wrong project's fact by id. The summary line (also used by mem review's bucketed listings) now appends the binding for non-global facts ([kind/status @scopeRoot]); global facts and --json output are unchanged.

  • file-absent affirmed through a symlink, certifying a present file as removed -- existsFile (src/anchors.ts) resolved a symlink refusal (containsSymlink, at the target or at any intermediate directory) to a plain false, and file-absent mapped that false straight to affirmed. Reproduced against the built bundle with a pnpm-style junction, node_modules/foo -> node_modules/.pnpm/foo, both inside the anchor root, with node_modules/.pnpm/foo/package.json present on disk: mem remember "dependency foo was removed" --anchor "file-absent node_modules/foo" showed freshness=affirmed on mem show, asserting as verified ground truth that the dependency was gone while it demonstrably was not. file-absent node_modules/foo/package.json affirmed the same way. Any project with a symlinked node_modules, src, or packages directory got this fabrication for every file-absent anchored beneath it -- and "we removed dependency X" anchored file-absent node_modules/X is precisely the anchor an agent writes. file-exists had the opposite-but-equally-wrong instinct: it mapped the same refusal to contradicted, asserting the file definitely does not exist, when the honest answer is that mem cannot see through the symlink either way.

    evaluateTokens now checks containsSymlink for both file-exists and file-absent before calling existsFile, and returns unverified for either predicate on a detected symlink escape, rather than letting a false false flow into existsFile's own boolean. existsFile no longer does its own symlink check at all -- by the time it runs, the caller has already ruled that case out. This is a straightforward P3 read: a symlink means mem cannot safely resolve the path, so it can assert neither presence nor absence, and contradicted/affirmed were each a lie for one half of the file-exists/file-absent pair. Two existing tests encoded the bug as intended (asserting contradicted/affirmed for the outside-root symlink case) and are flipped to unverified; a new inside-root case covers the exact pnpm junction shape above.

  • glob-exists contradicted patterns that named files which plainly existed -- three separate bugs stacked in the same function. First, evaluateGlobExists split a pattern only on /, so a leading ./ (./src/*.ts) left a stray . segment that never matches a real directory entry, and on Windows a pattern typed with backslashes (src\*.ts) was one opaque segment instead of two. Second, and worse, the walk refused to descend into a directory literally named .git or node_modules even when the pattern named it directly -- glob-exists node_modules/pkg/index.js and glob-exists node_modules/** both contradicted a file that existed, because the hardcoded skip did not distinguish "the wildcard happened to match this name" from "the pattern explicitly asked for this directory". Reproduced against the built bundle with src/a.ts and node_modules/pkg/index.js both present: glob node_modules/pkg/index.js, glob node_modules/**, and glob ./src/*.ts all recorded contradicted and were excluded from ground truth and flagged under mem review for the user to forget, while only the exact-syntax glob src/*.ts came back correct.

    Pattern splitting now also splits on \ when FS_CASE_INSENSITIVE (win32) is set, and drops empty . segments rather than choking on them. The .git/node_modules skip now applies only when the segment that matched the entry was itself a wildcard (*, ?, or the recursive segment): a literal segment naming .git or node_modules is treated as an explicit request to descend and is honored, while */pkg/target.txt still refuses to walk into node_modules reached only by the wildcard *. The existing budget (MAX_GLOB_ENTRIES_SCANNED, yielding unverified on overrun) st...

Read more

v0.3.1

Choose a tag to compare

@github-actions github-actions released this 31 Aug 03:32

Fixed

  • The agent-facing seam withheld facts while its output claimed to be complete -- when buildHintFormat exceeded its 150 ms soft budget it dropped its emission caps from 8/4 to 2/1 and returned the smaller set. Measured against a real store: 500 facts returned 12 lines, 2,000 returned 12, 10,000 returned 8, and 40,000 returned 2. The payload at 40,000 was byte-shape-identical to a healthy one -- same TGMEM/2 header, same line grammar, same footer -- so a consumer had no way to tell 2 facts from all of them. HintFormatResult.truncated was computed and returned, but src/cli.ts never read it, and the only other signal was a logWarning on stderr, which the documented consumer contract does not read: README describes exactly two consumer failure modes, a missing binary and a caller timeout, and this is neither. Truncation fires when mem overruns its own budget and still answers inside the caller's window, so the caller's fail-open path never engages. For a seam whose stated guarantee is self-caveating output, that is the one failure it is not allowed to have.

    There is no way to say "this is partial" in TGMEM/2. The grammar is closed: a conforming consumer drops an off-grammar line rather than guessing at it, footer-text is pinned to one exact string, and any grammar change bumps the version -- at which point consumers that have not upgraded fail open to no hints at all. So annotating a reduced response is not available, and a reduced response cannot be distinguished from a complete one. The budget-exhausted path now returns an empty hint set instead, which is already this module's shape for "I could not deliver" (buildHintFormat's internal-failure catch returns the same thing) and which the consumer's fail-open path already handles. TRUNCATED_AGGRESSIVE_CAP and TRUNCATED_PRECISION_CAP are gone; no wire change and no version bump were needed.

    Verified end to end against the built bundle at 40,001 facts over six runs: output is now strictly binary -- 10 lines (header, 8 fact-lines, footer) or 1 line (bare TGMEM/2), never a subset. Both regression tests were confirmed failing before the fix, emitting 4 and 2 lines respectively. The in-process test asserts lines is exactly empty rather than merely shorter than a healthy response, because a shorter-than assertion would pass again the moment a reduced cap was reintroduced, which is the defect; the wire-level test asserts no fact-lines and no footer reach stdout, since TGMEM/2 emits a footer only alongside at least one fact-line and a lone footer would itself be off-grammar.

  • The seam starved one whole fact kind at trivially small store sizes -- buildHintFormat splits results into an aggressive set (preferences and corrections, cap 8) and a precision set (everything else, cap 4), but it called retrieve() without a limit, so it received DEFAULT_RECALL_LIMIT (20) results and split those. That limit is a post-ranking slice applied before the kind-split, so whenever the top 20 shared one kind the other cap was starved to zero. It could never have bounded the output anyway -- the caps total 12, already under 20 -- so its only effect was to distort composition.

    Score ties break on captured_at descending, which makes the triggering shape an ordinary one rather than an extreme one: older decisions, newer preferences accumulating on top. Reproduced against the built bundle on a 33-fact store (8 decisions recorded first, then 25 preferences): the pre-fix binary emitted 8 preference lines and zero decision lines, the post-fix binary emits 8 and 4, from the same database and the same command. Every decision the user had recorded was invisible to their agent, in a payload carrying the normal header, grammar and footer -- indistinguishable from "this project has no decisions."

    This is the third path in the class the entry above describes, and the one that reaches a real user first: it needs no timing pressure and no unusual scale, only more recent facts of one kind than the recall limit. HINT_FORMAT_RECALL_LIMIT now makes the seam's own caps the sole bound on what reaches the wire. Found by the new scale-invariant test rather than by hand -- the assertion that 500 seeded facts emit exactly 8 + 4 failed at 8 + 0.

  • Nothing exercised the seam at a store size where its caps bind -- every existing test seeded a handful of facts, so the emission caps, the kind-split and the budget path were only ever tested where they could not interact. tests/unit/integration-seam.test.ts now seeds 500 facts and asserts the invariant directly: the fact-line count is the full cap set or zero, never a third size, with the wire-shape assertion that a complete set carries the footer and an empty one carries nothing. The third of those tests deliberately runs against the default budget so it exercises whichever path the runner takes, which is what makes it an invariant guard rather than a latency assertion -- recall is linear in store size, and a wall-clock assertion on a shared runner is the flake this file already carries a wrapper to prevent.

  • The budget-pinning convention was a convention, and one file had already missed it -- tests/unit/integration-seam.test.ts gained a wrapper defaulting the retrieval soft budget in 0.3.0; tests/integration-seam.test.ts was never given one. That stayed survivable only while budget exhaustion shrank the result, since a content assertion could still pass against the smaller set. Once exhaustion emptied the result instead, the same latent flake turned two Windows CI jobs red. tests/guards/seam-budget.test.ts now fails if any test file imports buildHintFormat without defining a pin, and carries a second assertion that at least one file is being watched, so a rename cannot quietly reduce it to a test that asserts an empty list is empty.

  • 26 tests asserted an error class where the class could not identify the failure -- expect(...).toThrow(SomeError) passes for any instance of that class from any code path, so wherever one class covers several distinct failures the assertion silently weakens into "something went wrong." WiringConflictError carries 14 distinct messages and CaptureValidationError 19; a test named for one guard was free to pass on any of the others. This was not hypothetical: 0.3.0's own oversized-import test, named "throws JsonImportError before attempting to parse," was passing on the parse -- 50MB of filler is not valid JSON either, and raising the size limit tenfold left the test green, meaning the guard it existed for had never been covered.

    Every one of the 26 now pins the message alongside the class. Auditing them found no further live defects -- each was throwing what its name claimed -- but four sites in tests/capture.test.ts and tests/unit/capture.test.ts were missed by the manual sweep and caught only by the new guard, which is the argument for having one. tests/guards/error-assertions.test.ts fails when a bare class assertion names a class src/ can raise with more than one message. Its ambiguity model counts construction sites rather than throw sites, because readFileWithErrorMapping (src/fileUtils.ts) constructs the caller's class across four message branches -- which is how JsonImportError reaches nine possible messages from five direct throws, and why MarkdownImportError, never thrown directly at all, would otherwise have counted as unambiguous. Classes defined inside a test file are exempt: tests/unit/fileUtils.test.ts passes TestError in to prove the mapper returns the class it was given, so there the class is the assertion.

    Verified in both directions. Stripping one message assertion makes the guard fail naming the file, class and message count -- while the 85-test wiring suite it came from stays entirely green, which is precisely the regression the guard exists to catch and the suite cannot see.

  • mem recall <query> returned the whole store for a query that matched nothing -- byte-identical to mem recall with no query at all. A query is a ranking input, not a filter: BM25 orders the candidate set and never removes from it, so the results.length === 0 branch that prints no matching facts could only ever fire when a --kind/--scope filter excluded everything, never when the query itself matched nothing. On a three-fact store, mem recall xyzzyplughquux returned all three facts, ordered by recency, presented exactly as a hit. The no matching facts outcome src/cli.ts documents was unreachable by the query path.

    Recall now says so and still shows the facts: the fix adds the missing signal rather than emptying the result. The condition is that every result scored 0, which is the documented meaning of an empty query ("all candidates tie at score 0", src/retrieval.ts) and also covers a term so common it appears in every fact -- zero discriminating power, so "did not narrow these results" is true there too, which is why the wording claims that rather than claiming the term is absent. The note is human-path only: --hint-format returns before it, so the closed TGMEM/2 grammar is untouched, and a regression test asserts no such line reaches the wire.

  • mem list --kind decision reported "no facts stored" on a store that was not empty -- the message is a claim about the whole store, and a filter excluding everything is a different fact about the world. A user filtering a populated store was told it held nothing. It now distinguishes the two, and still says "no facts stored" when the store really is empty.

  • exitOverride() covered the root command only, so all 15 subcommands exited mid-flush -- run() sets it so Commander cannot "call process.exit() mid-flush", per the comment that has always sat above the line. It is not inherited: a Command applies it to itself alone. Every subcommand therefore kept Commander's default behaviour and called process.exit() dire...

Read more

v0.3.0

Choose a tag to compare

@github-actions github-actions released this 30 Aug 03:31

Added

  • A tests/bundle/ tier that drives the built dist/token-goat-mem.mjs as a subprocess -- CONTRIBUTING.md's own rule is that a command with no coverage against the built bundle fails the gate, and until now exactly one test in the suite executed it (mem --version); everything else drove run() in-process against transformed TypeScript. The new tier is deliberately small: an init/uninstall round trip asserting hand-formatted config files come back byte-identical (four-space indent, CRLF, no trailing newline, no empty containers left behind), and a remember/recall/forget smoke test guarding what only a real subprocess can see -- externals resolving at runtime, the esbuild define, native-module loading, exit codes, and stdout vs stderr. Verified by dropping better-sqlite3 from esbuild's external list: 8 of the 9 new tests fail on that build while the in-process suite stays green.

  • The pre-commit gate CONTRIBUTING.md specifies is now a real hook -- npm run lint && npm run typecheck && npm run test:guards was documented as the check to run before every commit, with no .husky, no .git/hooks/pre-commit, and no hook manager behind it; it was a convention someone had to remember to type, and a commit that skipped it was indistinguishable from one that passed. .githooks/pre-commit now runs it, enabled by a prepare script that points core.hooksPath at the directory on npm install. It stays deliberately narrow -- the guards tier is pure introspection with no I/O, so the hook is fast enough that nobody reaches for --no-verify out of habit, and npm test stays a pre-push concern. Verified by staging a deliberate type error and confirming the commit was refused.

Fixed

  • mem recall | head -1 crashed instead of exiting -- a reader that closes the pipe early, which is exactly what head, grep -q, and any | less the user quits out of do, left the next write failing with EPIPE. Nothing listened for error on stdout, so node promoted it to an unhandled error event: a stack trace and exit code 1 for a pipeline that did what the user asked. Terminating early is the reader's prerogative, not an error for the writer, so both output streams now exit quietly on EPIPE. Found by the new CI workflow on its first run -- the Node 18 Linux job's own recall | grep -q check is what died. Windows does not raise EPIPE for a closed pipe, so no amount of local testing on the development machine could have surfaced it.
  • Two seam tests asserted on wall-clock timing they did not control -- buildHintFormat truncates its own output when it exceeds a 150 ms soft budget, dropping its caps from 8/4 to 2/1. A cold CI runner can spend longer than that just opening the database, so a test asserting which facts came back was really asserting how busy the machine was: the first Windows CI run reported truncation on an empty store and returned one of three facts where three were expected. Pinning the budget per-test only fixed the two tests that happened to go red -- the next CI run turned up a third, and forcing the override on reveals six truncation-sensitive tests in that file. Every call in the suite now routes through a wrapper that defaults retrievalBudgetMs (a test override alongside the existing now and dbPath ones) to a budget no machine can exceed, so a new test cannot acquire the flake by omission; the one test that is about truncation overrides it back down to 0. The truncated path also gains the first deterministic coverage it has ever had -- until now it was only reachable by being unlucky.
  • SECURITY.md described secret screening only by what it catches -- "refuses to store secrets" reads as a guarantee, and the actual boundary is narrower: screening fires on named credential formats, a credential word joined to its value by a :/= separator or sitting within ~32 characters of a 32+ character hex run, and standalone high-entropy tokens of 32 characters or more. So password = Xk9mP2vL8nQ4wR is refused while the staging password is Xk9mP2vL8nQ4wR is stored verbatim -- no separator, not hex, under the entropy floor. That floor is what keeps ordinary project facts from being refused, so it is an accepted position rather than an open defect, and SECURITY.md now says so. The three worked examples are asserted in tests/unit/capture.test.ts, so the documented boundary cannot drift permissive without a test failing.
  • README used three terms before explaining them -- "seam" appeared in the badge line and opening paragraph with no definition; contested (two facts disagree with each other, ambiguous winner) and contradicted (an anchor predicate tested reality and denied the fact) were used adjacently with no cue that they are different mechanisms; and "source reference" (the always-populated source_ref column) collided by name with the "sources" table, which is wired but not yet written to by any capture path. Each is now defined where it first appears.
  • Coverage was configured and could not run -- vitest.config.ts had named the v8 provider and its reporters since the suite was created, with no @vitest/coverage-v8 installed, no script to invoke it, and no thresholds, so vitest run --coverage failed on a missing dependency and nothing ever produced a report. The provider is now a devDependency, npm run test:coverage runs it scoped to src/, and CI runs that instead of a bare npm test. The floors are ratchets set a couple of points under what the suite actually reaches (92.5 / 85.0 / 98.8 / 92.3 at the time of writing), so a module losing its tests fails the build while ordinary branch-level noise does not. Verified by raising the statement floor to 99 and confirming the run fails.
  • src/fileUtils.ts had no test of its own -- it is the boundary where a raw errno becomes a message a user reads, and its branches were only ever reached incidentally through whichever importer happened to hit a missing file, so the messages themselves were unpinned. It now has a test file covering each mapped code and the fallback branch that no filesystem state can produce.
  • A mem import that overlapped another by one fact lost every fact it was going to add -- the duplicate-id check that decides whether to insert a fact ran in the classification pass, which is outside the transaction that takes the write lock. Two imports sharing one id could therefore both classify it as new; the loser then hit the facts.id primary key, and because that insert sits inside the batch transaction, the whole batch rolled back rather than the one row that was already there. The id is now re-checked under the write lock and recorded as an ordinary duplicate skip. Individual writers were always correctly serialized; it was the read backing the decision that sat outside the lock. The regression test reproduces the interleave exactly rather than racing two processes: db is proxied so the rival's insert lands on the first db.transaction(...) call, which is the line immediately after classification and immediately before the lock is taken.
  • Preservation tests could not see the formatting damage they existed to catch -- the no-op and conflict-abort tests in wiring.test.ts asserted toEqual on parsed objects, which is blind to exactly the failure mode 0.2.6 shipped fixes for: an indent silently normalized from four spaces to two, a compact file exploded across lines, a trailing newline added or dropped. They now assert on bytes, and their fixtures are seeded non-canonically (four-space indent, compact and unterminated) so a reserialization is visible rather than incidentally equal.
  • Type-aware linting was being paid for and not collected -- eslint.config.js set parserOptions.project for every .ts file, which is the expensive part, then extended the non-type-aware recommended preset. recommendedTypeChecked now runs on src/** (not tests/**, where the no-unsafe-* family fires ~57 times on fixtures that parse JSON into any and immediately assert on it -- a legitimate shape for a fixture). It found seven issues, all fixed: three type assertions the surrounding narrowing already made redundant, two async handlers with nothing to await (guard accepts void | Promise<void>, so the keyword bought nothing), and new Array(n).fill(undefined) inferring any[] and being silently assigned into a typed array in importFromJson. One rule is suppressed with its reasoning at the call site: withTimeout forwards a caller-supplied embedding backend's own rejection reason verbatim, and wrapping it in an Error would bury what the caller needs to diagnose it. The rule earning this on its own is no-floating-promises, which finds zero violations today -- verified biting on an unawaited database write.
  • Two runtime dependencies nobody used were installed on every machine -- zod@^4.4.3 sat in dependencies imported by zero lines of src/ and absent from the shipped bundle, and sqlite-vec@^0.1.9 sat in optionalDependencies referenced only by a comment, through six releases. Both are removed. sqlite-vec can return the day a concrete embedding backend actually loads it; nothing today can reach the EmbeddingBackend seam in retrieval.ts, so it was install weight for an unreachable path. Verified by installing the packed tarball into a clean prefix: neither package appears, and --version, doctor, and a remember/recall round trip all work.
  • CONTRIBUTING.md documented a runtime surface that was wrong in both directions -- it named zod, which was unused, and omitted jsonc-parser, which is real, marked external in the esbuild config, and resolved from node_modules at runtime. A new guard (tests/guards/dependencies.test.ts) now fails when a declared runtime dependency is imported nowhere in src/, so an unused one cannot sit in the manifest unnoticed again.
  • **Three integration docs told users to paste text `mem in...
Read more

v0.2.6

Choose a tag to compare

@github-actions github-actions released this 29 Aug 05:15

Every fix in this release is in mem init / mem uninstall. The documented promise for those two
commands is that uninstall reverses exactly what init wrote and leaves everything else alone; five
separate defects broke it, one of them by deleting the user's own text.

Fixed

  • mem uninstall could delete a block of the user's own file -- the per-tool marked block was located with a bare indexOf(start) / indexOf(end) pair, which pairs the first start marker with the first end marker even when they do not belong together. A hand-edit, an interrupted write, or a merge conflict can leave an orphaned <!-- token-goat-mem:<tool>:start --> behind with no end of its own; uninstall then paired that orphan with the end marker of the real block further down and removed every byte in between. In a CLAUDE.md shaped # My notes / orphan / user prose / real block, the user prose went with it. The block is now located by scanning every start marker and taking the first one that resolves to a complete pair, which is how the shared AGENTS.md block has always been located -- the per-tool path simply never got the same rule. A stray end marker sitting ahead of the real block no longer makes install append a second copy, either.
  • A reinstall never upgraded the shared AGENTS.md block's body -- codex, copilot-cli, and copilot-vscode share one reference-counted block, and the writer returned early as soon as the tool was already named in tools=. A body written by an older version of mem was therefore left in place forever: the per-tool blocks were replaced with current text on every reinstall and this one silently drifted away from them. It is now rebuilt from the same constant every time, and returns the file untouched only when the rebuild is byte-identical.
  • mem init rewrote the formatting of files it does not own -- .claude/settings.json was round-tripped through JSON.stringify(parsed, null, 2), so a config indented with four spaces came back indented with two, with its key layout and any non-canonical spacing gone. .vscode/tasks.json and keybindings.json went through jsonc-parser's modify with a formattingOptions, which reformats the entire containing array, so a hand-written one-line { "key": "ctrl+q", "command": "noop" } came back exploded across five lines. Neither file is mem's to restyle. Both paths now edit text rather than reserialize an object: modify with no formatting options yields a genuinely surgical edit that touches nothing else, and mem re-renders only its own inserted payload, at the document's own indent unit (tabs included) and line ending. Comments and trailing commas survive in both directions.
  • mem uninstall left behind the containers mem init had created -- installing into a settings.json with no hooks at all, which is the common case, created hooks and hooks.SessionStart; uninstall removed the hook group and stopped, leaving a "hooks": { "SessionStart": [] } husk. tasks.json had the same shape via "inputs": []. Uninstall now prunes an array or object that its own removals emptied. An empty hook array is inert, so pruning one a user happened to have written by hand costs them nothing.
  • mem init wrote LF into files authored with CRLF, and mem uninstall grew a blank line each time -- the marked block was assembled with hard-coded \n, so installing into a CRLF-authored CLAUDE.md produced a file with mixed line endings, and rewriting the shared block's tools= line dropped the \r from it. Uninstall then looked for a literal "\n\n" separator it would never find in such a file and left the surrounding blank lines in place, so each install/uninstall cycle added one. Every write now detects and preserves the file's own line ending, and the separator inserted by install is exactly the one removed by uninstall -- so a file with no trailing newline no longer gains one, and a whitespace-only file is no longer emptied.
  • The repository's own tracked files can no longer be checked out as CRLF -- a .gitattributes with * text=auto eol=lf pins them. Under core.autocrlf=true, which is the Windows default, docs/integrations/copilot-vscode.md arrived with CRLF on all 161 lines, and the doc/code consistency test added in 0.2.5 parses that file with a /```json\n/ regex that cannot match \r\n. That is a test that fails for a reason having nothing to do with what it tests, on a checkout that is otherwise correct.
  • The dry run described a shared-block install it was no longer doing -- now that a reinstall refreshes a stale body, a tool already named in tools= can have work to do, and mem init --dry-run still reported it as "install would join existing shared block (adds <tool> to tools=)". That case now says it would refresh the body.
  • The npm audit note in CONTRIBUTING.md described advisories that no longer exist -- it documented five dev-only advisories in the esbuild/vite/vitest chain and the reasoning for not forcing the fix. The toolchain upgrade that cleared all five landed before 0.2.5; npm audit has reported zero since, and the paragraph telling contributors otherwise outlived it.

Added

  • Nine regression tests that assert on file bytes, not parsed structure -- the formatting damage above was invisible to the existing suite because every assertion went through JSON.parse, and the round-trip leaks were worse than invisible: two tests had encoded the leftover empty containers as the expected result, one of them inside a test named "uninstall returns the file to its pre-install state". The new tests seed hand-formatted fixtures -- four-space, tab-indented, single-line, comment-carrying, CRLF -- and compare the file to itself after an install/uninstall cycle. Five were verified failing against the previous implementation before the fix landed.
  • Three regression tests for malformed and stale markers -- an orphaned start marker ahead of a real block, a stray end marker before one, and a shared block whose body was written by an older version. Each was verified failing first.
  • A ## Line endings section in CONTRIBUTING.md -- explains what .gitattributes pins, when to run git add --renormalize ., and the distinction that matters here: files mem edits are the user's, and code that writes them must never hard-code \n.

v0.2.5

Choose a tag to compare

@github-actions github-actions released this 28 Aug 17:20

Documentation and CI only -- no runtime change. mem behaves identically to 0.2.4.

Fixed

  • The README no longer contradicts itself about installation -- it carried two install sections, one saying npm install -g token-goat-mem and one saying "Not yet published to npm -- install from source". The claim was written at v0.1.0 when it was true; 0.2.0 updated the block at the top of the file and missed the ## Install section further down, so the file has told readers both things at once through every release since. The duplication was the cause, so it is gone: ## Install is now the single place carrying the npm command, the from-source steps, requirements and verification, and the top of the README is a one-line quick start that links to it. Published packages snapshot the README at publish time, so npmjs.com showed the stale text until this release.
  • The copilot-vscode integration doc documented the keybindings 0.2.4 replaced -- that release moved ctrl+shift+m/ctrl+shift+n to the chords ctrl+k m/ctrl+k r because the originals shadowed View: Problems and New Window, but only src/wiring.ts was updated. The doc presents itself as "what mem init copilot-vscode writes, if you'd rather do it by hand", so anyone following it installed by hand the exact two shadowing bindings the release had just removed.

Added

  • A test asserting that integration doc matches what the installer actually writes -- the keybinding drift above shipped with a green suite, because the regression test for it compared code against code and nothing looked at the documentation. The new test installs into a temp fixture and deep-compares the result against the fenced JSON blocks in the doc, covering both the keybindings and the tasks/input block, so the doc's own promise is the assertion rather than a third restatement of the values. Verified against the real drift: with the doc reverted to its 0.2.4 state, it fails with the exact ctrl+shift+m to ctrl+k m diff.

Changed

  • CI moved actions/checkout and actions/setup-node from v4 to v7 -- the v4 actions target Node 20, which the runners now force onto Node 24. The breaking changes across the two intervening majors were checked against this workflow: setup-node v5's automatic caching activates only on a packageManager field, which this package does not have, and registry-url/NODE_AUTH_TOKEN (which the publish step depends on) are unchanged throughout. node-version: 20 is untouched -- it selects the Node that builds the package, not the action runtime the deprecation concerns.