Skip to content

v0.5.0

Choose a tag to compare

@github-actions github-actions released this 02 Oct 04:24
· 64 commits to master since this release

Added

  • On Windows, mem init claude-code --user writes hooks that launch node on the installed bundle directly. npm's mem shim is a shell wrapper that adds a noticeable delay to every hook event, so the user-level hook command is now if [ -f "<bundle>" ]; then node "<bundle>" <subcommand> ... || <fallback>; elif command -v mem >/dev/null 2>&1; then mem <subcommand> ... || <fallback>; fi, falling back to the plain mem guard when the bundle is gone. The bundle path comes from the running install, uses forward slashes, and is refused (plain shape kept) if it contains ", $ or a backtick. Project hooks and non-Windows installs keep their exact commands. Re-running init upgrades an old-shape hook in place without duplicating it, uninstall removes either shape, and the PATH pre-flight probes the bundle itself when it is the launcher. mem doctor compares project and user hooks by subcommand and flags, so a launcher-only difference is no longer reported as divergence, and a new launch-time finding warns when mem --version takes over 300 ms, measured only when some installed hook would launch the PATH mem (on Windows it suggests re-running mem init claude-code --user).

  • mem review --promote <id> --subject <key> --value <value> keys a pending fact as it activates it. Keying, activation, and the contradiction pass the new key makes the fact eligible for run in one transaction, and the audit row records keyed <subject>=<value>, so a suggested fact no longer needs a follow-up mem edit to become supersedable. Both flags are required together, only with --promote, and are validated and secret-screened like mem remember's; a contested fact is already keyed and is refused with a pointer to mem edit.

  • mem epoch --json emits the current write epoch as JSON. Outputs { epoch: <number> }. The plain output remains unchanged byte-for-byte for cache invalidation.

  • mem facets --json emits structured facets data for each mode. --list-entities --json emits an array of { term, facts } objects; --fact <id> --json emits { fact, entities, topics }; backfill/--all mode emits { facts, entities, topics } summary. Plain output is unchanged.

  • mem backup --list --json emits snapshot records as JSON. Outputs an array of snapshots, each with path, name, takenAt (ISO 8601), epoch, reason, and size fields. The plain output remains unchanged.

  • mem doctor checks wiring drift for every supported tool, and project-vs-user hook divergence. For each tool in mem init's list whose config exists under --root (new on doctor, default the current directory) or the home directory, a wiring finding is ok when the block matches exactly what mem init <tool> would write now, and warn (remedy mem init <tool>, with --user at user level) when it is outdated, hand-edited inside its markers, or a Claude Code hook lacks an expected event. Tools mem was never wired into are one informational ok line, and doctor still writes nothing. A new hook-divergence finding warns when project and user Claude Code hooks run different mem invocations (both run, so recall is duplicated or inconsistent) and names which to remove or re-init.

  • mem forget <id> --by <winner-id> records which fact replaced the forgotten one. Plain forget retires a fact with no named successor, so mem show and mem export could only say superseded_by: unknown. --by takes the winner by id or unique prefix (same resolver and errors as the id argument), refuses the fact itself or a winner that is not active/pinned, and writes the same Superseded by fact <id> audit detail the contradiction and review paths use, in the same transaction as the forget. mem review's pending may contradict line now ends with a paste-ready mem forget <rival> --by <id>.

  • Retrieval now indexes a decision's rationale for lexical search. computeBm25Scores now appends doc.why to the tokenized document text, so queries matching only a decision's rationale (e.g., "staging memory") can find the decision that records it. This lets "why did we..." questions in a future session surface decisions before relitigating them, without waiting for embeddings to refresh.

  • mem doctor --json and --strict. Every doctor line is now a structured finding
    ({ check, status: ok|warn|fail, message, remedy? }) drawn from a fixed list of check names, and
    the plain report is those messages printed unchanged. --json prints { findings, epoch }.
    --strict exits 1 when any finding is fail (an installed hook whose mem binary is missing or
    too old to run it); without it doctor still exits 0, so a warning such as a store with facts and no
    backup never breaks a script that did not ask.

  • npm i -g token-goat-mem installs the Claude Code hooks. A new postinstall script runs the freshly installed mem init claude-code --user on every global install and upgrade, so recall reaches every project without a manual step and an upgrade refreshes hooks an older mem wrote. It keeps mem init's pre-flight check that the mem on PATH can run those hooks, never fails the install (anything short of success prints mem init claude-code --user as the fix), does nothing on a local install, and is skipped with TOKEN_GOAT_MEM_SKIP_HOOKS=1. It also leaves the hooks alone when run as someone other than the home directory's owner (root under sudo npm i -g, which would otherwise leave a root-owned ~/.claude/settings.json), and stops an init that hangs after two minutes (TOKEN_GOAT_MEM_POSTINSTALL_TIMEOUT_MS). npm is the only package manager that runs it: pnpm, yarn, bun and --ignore-scripts installs need mem init claude-code --user by hand.

  • The store is backed up outside the mem home. Snapshots land in ~/.mem-backups (override: TOKEN_GOAT_MEM_BACKUP_DIR), so deleting ~/.mem -- or ~/.claude, which never held the store -- no longer loses every fact. Opening an existing store takes an automatic VACUUM INTO snapshot at most once a day and only when it has changed (the newest 14 are kept), and always before a pending schema migration (once per store state, so a migration that keeps failing does not copy the store on every open). Only the open that wins a claim file in the backup directory takes the daily snapshot, so hooks opening the store from several processes at once copy it once; a claim or partial copy left by a process that died is cleared after an hour. Each snapshot is named for the epoch of the copy itself, never a read taken before it. mem backup takes one by hand and mem backup --list lists them; mem restore <snapshot> checks and migrates a private copy of the snapshot first -- it must pass integrity_check, hold mem's own facts columns, and not come from a newer mem -- then saves the store it replaces as a pre-restore snapshot and swaps the rows in one transaction on the live connection, advancing the epoch past both stores. A write that lands between that snapshot and the swap is caught and the swap retried, so the pre-restore snapshot always holds exactly what was replaced, and the restore never prunes the automatic snapshot it is restoring from. A snapshot that cannot be written never fails the command that opened the store; mem doctor gains a backups: line that reports it instead.

  • mem scan-session can file a correction. Its trigger table yielded fact, preference and
    decision and had no shape that produced a correction, while the block mem installs tells the
    agent to persist exactly those three kinds plus corrections. The scanner was structurally unable to
    find one. Two openers now do: an explicit correction: prefix, matching the existing decision:
    and rule: entries, and the reversal forms of that's wrong / that's outdated / that's no longer true. Bare no, and actually, are deliberately not among them. They open any negative
    answer -- "no, that test is fine" reverses nothing -- and the pending queue is sorted by sighting
    count, so a false positive restated across sessions would climb it. Only shapes that assert the
    reversal in their own words qualify; anything subtler is what mem remember --kind correction is
    for.

  • A cross-process concurrency test. mem has no daemon: every command is a short-lived process
    contending for one WAL database under BEGIN IMMEDIATE, and no in-process test can exercise that,
    because a single Database handle serialises everything by construction.
    tests/bundle/concurrency.test.ts runs six workers of the built bundle against one store that does
    not exist yet, so schema creation races too. Each worker interleaves unique remembers, a
    remember of one shared sentence, and recalls. The test asserts no write is lost, the shared
    sentence stays one row with every later statement audited as a reaffirm, and nothing reports
    SQLITE_BUSY. Verified by opening the database with a zero busy timeout: recall's surfaced-marking
    write then reports a locked database, and the test fails.

  • mem log: the store-wide audit timeline. mem show <id> could read one fact's audit trail,
    but only once you knew which fact to ask about; "what did the agent change in my memory this
    week?" had no answer. mem log lists every audit row newest first, each prefixed with a pasteable
    short fact id, with --fact, --event (exact name or capture-style family), --age-days,
    --limit, and --json. --fact also resolves ids gc has hard-deleted, because audit rows
    are kept for 180 days and superseded facts for 90, and it reports a prefix shared by a live and a
    deleted fact as ambiguous instead of showing only the live one. mem show's history block now
    renders through the same formatter.

  • mem reflect: resolve pending suggestions while the agent still knows what it meant.
    mem scan-session files every durable-sounding sentence as pending and stops there, so the
    queue only grew, and the one party that knew whether "never commit generated files" restated,
    changed, or added to a stored fact -- the agent that said it -- had moved on. mem reflect lists
    each pending suggestion beside the live facts it most resembles, with the resolutions in the order
    to try them: update the related fact first, promote as new only when nothing covers it, reject
    otherwise. --transcript files a transcript and lists only its suggestions; --hook-stdin is the
    new Stop hook (below).

  • mem init opencode. opencode reads AGENTS.md, so a project install joins the shared,
    reference-counted "## Memory" block codex, copilot-cli and copilot-vscode already write there.
    Unlike those three, opencode also has a global rules file, so --user is supported: it writes
    ~/.config/opencode/AGENTS.md. That path is the same on Windows, because opencode resolves its
    config directory through xdg-basedir, which has no %APPDATA% branch. See
    docs/integrations/opencode.md.

  • A fact can carry its reason. A stored decision without its rationale is the one a later
    session is most tempted to relitigate: "the build uses esbuild" invites "why not tsc?", where the
    reason answers it before it is asked. mem remember/mem suggest --why "<reason>" records one
    (optional, at most 500 characters, secret-screened like every other stored field), mem edit --why changes it and --why "" clears it (both reversible with --undo), and restating a fact
    with a new reason replaces the stored one and logs why updated. The reason rides on the recalled
    display line as (why: ...) so the agent reading recall sees it too, except in terse mode, and
    survives mem export/mem import --from-json. The installed "## Memory" blocks now ask for a
    --why on decisions and corrections. Stored in a new nullable why column (schema v6); every
    existing fact reads as having no recorded reason.

  • A review decision can carry its reason. mem review --promote/--reject/--undo <id> --reason "<text>" appends ; reason: <text> to that transition's audit row (and, for a contested
    promotion, to each rival's supersession row), so mem log --fact <id> answers "why was this
    rejected?" later. Validated like --why (non-empty, at most 500 characters) and secret-screened
    against the same .mem/allowlist, since the audit log is as durable as the facts (a refusal is
    logged as review_blocked_secret, naming the pattern, never the value); --reason with no action
    to attach it to is refused.

  • mem doctor reports why coverage. A new line counts the active or pinned decisions and
    corrections that carry a --why (why coverage: 3/5 ...) and, while any lack one, names
    mem edit <id> --why "<reason>" as the fix. Pending and superseded facts are left out: recall
    never surfaces them, so their missing reason costs nothing.

Changed

  • A more specific same-subject fact now overrides a broader one at recall. Contradiction
    resolution keys on subject + scope, so a global package-manager = npm and a project
    package-manager = pnpm both stayed active and recall surfaced both, which is conflicting
    guidance. retrieve() (shared by mem recall and the TGMEM/2 hint seam) now withholds the
    broader fact from that recall when an in-scope narrower fact (path over project over global)
    has the same subject and a different value. The broader fact is not superseded or edited, so it
    still surfaces in other projects; a pin does not exempt it, and subject-less facts, equal
    values, and contested/pending facts are unaffected. A narrower fact whose anchor is
    contradicted is withheld itself, so it does not hide the broader one.

  • Releases publish only after the full CI matrix passes, with npm provenance. The publish
    workflow now calls the CI workflow and waits for it, so a tag on a commit that fails lint,
    typecheck, or a test on either OS no longer reaches npm. The package is published with
    --provenance, which links each release on npm to the commit and workflow run that built it.

  • The Claude Code Stop hook runs mem reflect instead of mem scan-session --quiet. It files
    the transcript exactly as before, then blocks the stop with the mem reflect worklist as the
    reason -- but only for suggestions that run filed, so a second stop over the same transcript is
    silent, and never while stop_hook_active is set. PreCompact keeps scan-session --quiet: a
    compaction has no agent turn to answer a worklist. Re-running mem init claude-code adopts an
    unstamped scan-session Stop hook from an older install and rewrites it in place rather than
    leaving both to run; mem doctor names reflect as the one subcommand an older PATH binary lacks.

Fixed

  • mem doctor inspects the store read-only and keeps reporting when it is broken. It used to open the store through the same path every write command does, so a health check created a missing database, ran pending migrations, took snapshots, and died with an exception on a corrupt file -- the one case it exists for. It now opens the file read-only (openDbReadOnly), runs the hooks, backups, embedding and dream checks regardless, and reports store: none (no file; nothing is created), store: unreadable (<code>) (a fail, so --strict exits 1), schema: N migrations pending, and integrity from PRAGMA quick_check as findings. --json's epoch is null when a store exists but could not be read, and such a store reports only the embedding endpoint, never vector counts it did not read.

  • mem doctor no longer reports does not support ? for the earliest hook shape. The first mem init claude-code stamped a bare, unguarded mem recall --hint-format --root "$CLAUDE_PROJECT_DIR"; the flag parser only knew the two guarded wrappers, so doctor read that hook as unparseable and printed a placeholder subcommand. Recognising a hook as mem's and reading its flags now share one list of wrapper shapes, so they cannot disagree, and a stamped command in no known shape names its remedy (re-run mem init claude-code) instead of ?.

  • mem scan-session no longer blocks a Windows session's suggestions as a "secret". It stamps
    each suggestion's sourceRef with the native transcript path, and Claude Code names a transcript
    after its session UUID. The generic high-entropy check exempts path-shaped tokens by their /,
    but a Windows \ split the path into bare segments, so any session whose UUID scored above the
    entropy cutoff (about 40% of them) had every suggestion rejected as
    capture_suggested_blocked_secret. A \ in sourceRef and anchor now counts as a path
    separator for that check; named secret patterns still run against the raw value.

  • Audit details no longer read "stored active fact fact". The capture confirmations stopped
    doubling the noun for --kind fact in an earlier release, but the audit details capture writes
    (stored active, stored pending, restated) kept their own <kind> fact templates, so
    mem show and mem log still printed "fact fact". Both now render through one shared noun helper.

  • A rejected suggestion came back as pending once the garbage collector pruned its tombstone.
    mem review --reject marks a fact superseded, and the retention pass hard-deletes superseded rows
    past 90 days or 1000 rows. Nothing distinguished a rejection from an ordinary contradiction loser.
    But the row is what keeps the sentence out of the queue: mem scan-session and mem import --from-md both dedup against stored text at any status, so deleting the tombstone deletes the only
    record that the question was ever asked. A stable CLAUDE.md re-imported each quarter refilled the
    queue with bullets a human had already declined, and the README's claim that rejected facts are
    "kept for audit" was false at the 91st day. A superseded fact whose prior status was pending is
    now exempt from both bounds, and does not consume a row-cap slot on the way past. The exemption is
    deliberately narrow: an ordinary loser, and anything reaching superseded from active, is pruned
    exactly as before. The honest cost is that rejection tombstones now grow without an upper bound;
    they are bounded in practice by how often a human rejects something, which is not often, and the
    mem epoch --gc documentation says so rather than leaving it to be discovered.

  • A value that agreed except in capitalization superseded the fact it agreed with. subject is
    stored through normalizeSubject, value was stored and compared raw, and contradiction detection
    keys on both. Two facts on one subject recording pnpm and Pnpm therefore read as a disagreement
    about the same question, and the recall-time gate resolved it by withholding the loser. The damage
    was invisible from mem show, which reported every row active and healthy, while recall quietly
    surfaced one of them. Comparison now runs through a normalizeValue that mirrors normalizeSubject
    -- trim, collapse internal whitespace, lowercase -- in both contradiction detection and the reaffirm
    match. The raw value is still what gets stored and printed, because what the user typed is what they
    should see. Genuinely different values still contradict; a default_branch of main against
    master is unaffected, which is the case the normalization must not swallow.

  • Two of the recall seam's three failure modes were byte-identical to having nothing to say. An
    unreadable store already emitted a footer saying so. A retrieval that exhausted its time budget, and
    any other internal fault, both returned an empty line set -- exactly the bare TGMEM/2 header a
    project with no memory yet emits. The hook commands end in || true and the warnings go to stderr,
    so nothing downstream distinguished them either: a slow disk on UserPromptSubmit told the agent
    this project had no memory while the store sat full and healthy. Each now carries its own footer
    clause. The budget clause says the hint set is empty because retrieval ran out of time rather than
    out of facts, and points at a plain mem recall; it promises no partial result, because a partial
    response is byte-indistinguishable from a complete one on this wire and emitting one would hand a
    consumer a subset its own contract says is everything. Neither clause needs a protocol bump: footer
    text and footer presence have always been outside that set.

  • mem review listed two buckets of facts and named no command that would accept them.
    --promote and --reject take pending and contested; the anchor-contradicted and overdue-pin
    buckets hold active facts, so the two verbs their sibling buckets advertise exit 1 on everything
    listed there. Recall's own caveat pointed at mem review to resolve a contradicted fact, which
    delivered the user to exactly that dead end. Both buckets now print the commands that do work --
    mem forget for a fact that is no longer true, mem edit --anchor for one whose anchor is wrong,
    re-running mem pin to re-confirm a stale pin -- and the mem edit line carries --force when the
    fact is user-stated, because mem edit refuses those without it and a remedy that refuses is the
    defect this repeats. This is the same fix one release earlier applied to the contested bucket.

  • A decaying preference reported full confidence everywhere a human could look. A non-pinned
    preference decays on a 180-day half-life and stops being ground truth below 0.5, but mem show
    printed the stored confidence of 1 with no hint that recall had already demoted it, and the garbage
    collector's preferences_decayed_below_floor count named no ids and no remedy. mem show now
    prints the effective value beside the stored one and, below the floor, the two things that fix it:
    restating the preference, or pinning it. Healthy preferences print nothing extra -- decay is
    continuous, so a note gated on any decay at all would appear on every preference seconds after
    capture and mean nothing.

  • The integration block mem installs advertised six of its eleven anchor predicates. The agent
    reading that block is the one writing anchors, and it was never told file-contains,
    file-not-contains, git-branch-is, package-version or valid-until exist. valid-until is the
    natural anchor for the corrections the same block asks it to record, so the omission cost exactly
    the facts most likely to expire. All eleven are now listed in both installed block bodies, in the
    per-tool integration guides, and in the --anchor help on remember, suggest and edit, which
    previously named none at all.