Skip to content

v0.3.0

Choose a tag to compare

@github-actions github-actions released this 30 Aug 03:31
· 188 commits to master since this release

Added

  • A tests/bundle/ tier that drives the built dist/token-goat-mem.mjs as a subprocess -- CONTRIBUTING.md's own rule is that a command with no coverage against the built bundle fails the gate, and until now exactly one test in the suite executed it (mem --version); everything else drove run() in-process against transformed TypeScript. The new tier is deliberately small: an init/uninstall round trip asserting hand-formatted config files come back byte-identical (four-space indent, CRLF, no trailing newline, no empty containers left behind), and a remember/recall/forget smoke test guarding what only a real subprocess can see -- externals resolving at runtime, the esbuild define, native-module loading, exit codes, and stdout vs stderr. Verified by dropping better-sqlite3 from esbuild's external list: 8 of the 9 new tests fail on that build while the in-process suite stays green.

  • The pre-commit gate CONTRIBUTING.md specifies is now a real hook -- npm run lint && npm run typecheck && npm run test:guards was documented as the check to run before every commit, with no .husky, no .git/hooks/pre-commit, and no hook manager behind it; it was a convention someone had to remember to type, and a commit that skipped it was indistinguishable from one that passed. .githooks/pre-commit now runs it, enabled by a prepare script that points core.hooksPath at the directory on npm install. It stays deliberately narrow -- the guards tier is pure introspection with no I/O, so the hook is fast enough that nobody reaches for --no-verify out of habit, and npm test stays a pre-push concern. Verified by staging a deliberate type error and confirming the commit was refused.

Fixed

  • mem recall | head -1 crashed instead of exiting -- a reader that closes the pipe early, which is exactly what head, grep -q, and any | less the user quits out of do, left the next write failing with EPIPE. Nothing listened for error on stdout, so node promoted it to an unhandled error event: a stack trace and exit code 1 for a pipeline that did what the user asked. Terminating early is the reader's prerogative, not an error for the writer, so both output streams now exit quietly on EPIPE. Found by the new CI workflow on its first run -- the Node 18 Linux job's own recall | grep -q check is what died. Windows does not raise EPIPE for a closed pipe, so no amount of local testing on the development machine could have surfaced it.
  • Two seam tests asserted on wall-clock timing they did not control -- buildHintFormat truncates its own output when it exceeds a 150 ms soft budget, dropping its caps from 8/4 to 2/1. A cold CI runner can spend longer than that just opening the database, so a test asserting which facts came back was really asserting how busy the machine was: the first Windows CI run reported truncation on an empty store and returned one of three facts where three were expected. Pinning the budget per-test only fixed the two tests that happened to go red -- the next CI run turned up a third, and forcing the override on reveals six truncation-sensitive tests in that file. Every call in the suite now routes through a wrapper that defaults retrievalBudgetMs (a test override alongside the existing now and dbPath ones) to a budget no machine can exceed, so a new test cannot acquire the flake by omission; the one test that is about truncation overrides it back down to 0. The truncated path also gains the first deterministic coverage it has ever had -- until now it was only reachable by being unlucky.
  • SECURITY.md described secret screening only by what it catches -- "refuses to store secrets" reads as a guarantee, and the actual boundary is narrower: screening fires on named credential formats, a credential word joined to its value by a :/= separator or sitting within ~32 characters of a 32+ character hex run, and standalone high-entropy tokens of 32 characters or more. So password = Xk9mP2vL8nQ4wR is refused while the staging password is Xk9mP2vL8nQ4wR is stored verbatim -- no separator, not hex, under the entropy floor. That floor is what keeps ordinary project facts from being refused, so it is an accepted position rather than an open defect, and SECURITY.md now says so. The three worked examples are asserted in tests/unit/capture.test.ts, so the documented boundary cannot drift permissive without a test failing.
  • README used three terms before explaining them -- "seam" appeared in the badge line and opening paragraph with no definition; contested (two facts disagree with each other, ambiguous winner) and contradicted (an anchor predicate tested reality and denied the fact) were used adjacently with no cue that they are different mechanisms; and "source reference" (the always-populated source_ref column) collided by name with the "sources" table, which is wired but not yet written to by any capture path. Each is now defined where it first appears.
  • Coverage was configured and could not run -- vitest.config.ts had named the v8 provider and its reporters since the suite was created, with no @vitest/coverage-v8 installed, no script to invoke it, and no thresholds, so vitest run --coverage failed on a missing dependency and nothing ever produced a report. The provider is now a devDependency, npm run test:coverage runs it scoped to src/, and CI runs that instead of a bare npm test. The floors are ratchets set a couple of points under what the suite actually reaches (92.5 / 85.0 / 98.8 / 92.3 at the time of writing), so a module losing its tests fails the build while ordinary branch-level noise does not. Verified by raising the statement floor to 99 and confirming the run fails.
  • src/fileUtils.ts had no test of its own -- it is the boundary where a raw errno becomes a message a user reads, and its branches were only ever reached incidentally through whichever importer happened to hit a missing file, so the messages themselves were unpinned. It now has a test file covering each mapped code and the fallback branch that no filesystem state can produce.
  • A mem import that overlapped another by one fact lost every fact it was going to add -- the duplicate-id check that decides whether to insert a fact ran in the classification pass, which is outside the transaction that takes the write lock. Two imports sharing one id could therefore both classify it as new; the loser then hit the facts.id primary key, and because that insert sits inside the batch transaction, the whole batch rolled back rather than the one row that was already there. The id is now re-checked under the write lock and recorded as an ordinary duplicate skip. Individual writers were always correctly serialized; it was the read backing the decision that sat outside the lock. The regression test reproduces the interleave exactly rather than racing two processes: db is proxied so the rival's insert lands on the first db.transaction(...) call, which is the line immediately after classification and immediately before the lock is taken.
  • Preservation tests could not see the formatting damage they existed to catch -- the no-op and conflict-abort tests in wiring.test.ts asserted toEqual on parsed objects, which is blind to exactly the failure mode 0.2.6 shipped fixes for: an indent silently normalized from four spaces to two, a compact file exploded across lines, a trailing newline added or dropped. They now assert on bytes, and their fixtures are seeded non-canonically (four-space indent, compact and unterminated) so a reserialization is visible rather than incidentally equal.
  • Type-aware linting was being paid for and not collected -- eslint.config.js set parserOptions.project for every .ts file, which is the expensive part, then extended the non-type-aware recommended preset. recommendedTypeChecked now runs on src/** (not tests/**, where the no-unsafe-* family fires ~57 times on fixtures that parse JSON into any and immediately assert on it -- a legitimate shape for a fixture). It found seven issues, all fixed: three type assertions the surrounding narrowing already made redundant, two async handlers with nothing to await (guard accepts void | Promise<void>, so the keyword bought nothing), and new Array(n).fill(undefined) inferring any[] and being silently assigned into a typed array in importFromJson. One rule is suppressed with its reasoning at the call site: withTimeout forwards a caller-supplied embedding backend's own rejection reason verbatim, and wrapping it in an Error would bury what the caller needs to diagnose it. The rule earning this on its own is no-floating-promises, which finds zero violations today -- verified biting on an unawaited database write.
  • Two runtime dependencies nobody used were installed on every machine -- zod@^4.4.3 sat in dependencies imported by zero lines of src/ and absent from the shipped bundle, and sqlite-vec@^0.1.9 sat in optionalDependencies referenced only by a comment, through six releases. Both are removed. sqlite-vec can return the day a concrete embedding backend actually loads it; nothing today can reach the EmbeddingBackend seam in retrieval.ts, so it was install weight for an unreachable path. Verified by installing the packed tarball into a clean prefix: neither package appears, and --version, doctor, and a remember/recall round trip all work.
  • CONTRIBUTING.md documented a runtime surface that was wrong in both directions -- it named zod, which was unused, and omitted jsonc-parser, which is real, marked external in the esbuild config, and resolved from node_modules at runtime. A new guard (tests/guards/dependencies.test.ts) now fails when a declared runtime dependency is imported nowhere in src/, so an unused one cannot sit in the manifest unnoticed again.
  • Three integration docs told users to paste text mem init does not write -- the doc/code consistency test added in 0.2.5 only ever matched ```json fences, so every markdown block mem writes went unchecked and drifted. claude-code.md appended an (e.g. --subject package-manager --value pnpm) example the installer never writes, directly under a line promising to show exactly what it writes; all three AGENTS.md docs opened with "This machine has token-goat-mem installed" where the installer writes "token-goat-mem is installed", and carried a parenthetical list of trust caveats the block does not contain. The check now compares each doc's markdown fence against the block install() actually wrote to disk, rather than against a source constant, so a defect in the marker wrapping cannot slip past it either.
  • docs/integrations/copilot-vscode.md contradicted itself about its own keybindings -- its keybindings section correctly records the 0.2.4 move to the Ctrl+K M / Ctrl+K R chords and explains that the old pair was dropped for shadowing View: Problems and New Window, while its workflow walkthrough four sections later still told the reader to press Ctrl+Shift+M and Ctrl+Shift+N. The walkthrough is prose outside every fenced block, so neither the JSON nor the markdown consistency check could see it; a third check now asserts it names the keys the installer writes and not the superseded pair.
  • Nothing ran the checks on push, and nothing ran them on Windows -- .github/workflows/ held one file, triggered only by a v* tag or manual dispatch, whose steps were npm ci then npm publish; lint, typecheck, test and build ran only as a side effect of prepublishOnly, so a commit breaking any of them landed on master and surfaced at the next release, reported as "publish failed" without saying which check broke. The single job was ubuntu-latest only, while src/ carries six explicit process.platform === "win32" branches and every platform-gated test in the suite is skipped on one platform or the other -- the four db-permissions tests never ran on the machine this project is developed on, and the one Windows-only wiring test never ran in CI. A new ci.yml runs lint, typecheck and the full suite as separate named steps on both platforms for every push and pull request.
  • engines: { node: ">=18" } was never exercised -- and could not be, since vitest 4 requires Node 20 or newer, so the suite cannot run on 18 at all. That claim is about the shipped bundle, not the toolchain, so a second CI job builds the bundle on Node 18 on both platforms and drives it directly with no test runner: version, doctor, and a remember/recall round trip. It is the narrowest check that would catch a Node-20-only API reaching the artifact users install.
  • CONTRIBUTING.md described end-to-end coverage the suite did not have -- it told contributors npm test "includes end-to-end tests that build and exercise the shipped dist/token-goat-mem.mjs bundle," which was true of a single test asserting --version. It now describes the three tiers that actually exist and says which one a new test belongs in.
  • The 0.2.6 note over-claimed comment preservation -- "Comments and trailing commas survive in both directions" covered a paragraph spanning both config paths, but only .vscode/*.json is parsed as JSONC. .claude/settings.json is parsed as strict JSON, matching Claude Code's own format, and a commented one is refused with a conflict error rather than modified. The claim is now scoped to the path it holds for.

Changed

  • mem recall now binds facts to the project they were captured in -- --root reached only anchor evaluation, so the store was searched whole: standing in one project and running mem recall surfaced another project's decisions as if they were local. --scope project did not help, because it matched the scope label rather than the binding, narrowing to "scoped to some project" instead of "scoped to this one" -- so even the README's own mem recall --root . --scope project example returned the wrong project's facts. A project fact now surfaces only from its own root, a path fact only from a root containing its file, and global facts from anywhere, unchanged. The filter runs inside retrieve alongside every other filter rather than narrowing the SQL query, because a pre-filtered pool hides a fact's rival from contradiction resolution, whose reinstatement pass then reads that absence as "nothing contests this" -- and scope binding is the filter most likely to separate two rivals, since a contradiction is keyed on subject + scope. mem recall --hint-format already scoped this way; the two paths now agree.