Skip to content

v0.3.1

Choose a tag to compare

@github-actions github-actions released this 31 Aug 03:32
· 179 commits to master since this release

Fixed

  • The agent-facing seam withheld facts while its output claimed to be complete -- when buildHintFormat exceeded its 150 ms soft budget it dropped its emission caps from 8/4 to 2/1 and returned the smaller set. Measured against a real store: 500 facts returned 12 lines, 2,000 returned 12, 10,000 returned 8, and 40,000 returned 2. The payload at 40,000 was byte-shape-identical to a healthy one -- same TGMEM/2 header, same line grammar, same footer -- so a consumer had no way to tell 2 facts from all of them. HintFormatResult.truncated was computed and returned, but src/cli.ts never read it, and the only other signal was a logWarning on stderr, which the documented consumer contract does not read: README describes exactly two consumer failure modes, a missing binary and a caller timeout, and this is neither. Truncation fires when mem overruns its own budget and still answers inside the caller's window, so the caller's fail-open path never engages. For a seam whose stated guarantee is self-caveating output, that is the one failure it is not allowed to have.

    There is no way to say "this is partial" in TGMEM/2. The grammar is closed: a conforming consumer drops an off-grammar line rather than guessing at it, footer-text is pinned to one exact string, and any grammar change bumps the version -- at which point consumers that have not upgraded fail open to no hints at all. So annotating a reduced response is not available, and a reduced response cannot be distinguished from a complete one. The budget-exhausted path now returns an empty hint set instead, which is already this module's shape for "I could not deliver" (buildHintFormat's internal-failure catch returns the same thing) and which the consumer's fail-open path already handles. TRUNCATED_AGGRESSIVE_CAP and TRUNCATED_PRECISION_CAP are gone; no wire change and no version bump were needed.

    Verified end to end against the built bundle at 40,001 facts over six runs: output is now strictly binary -- 10 lines (header, 8 fact-lines, footer) or 1 line (bare TGMEM/2), never a subset. Both regression tests were confirmed failing before the fix, emitting 4 and 2 lines respectively. The in-process test asserts lines is exactly empty rather than merely shorter than a healthy response, because a shorter-than assertion would pass again the moment a reduced cap was reintroduced, which is the defect; the wire-level test asserts no fact-lines and no footer reach stdout, since TGMEM/2 emits a footer only alongside at least one fact-line and a lone footer would itself be off-grammar.

  • The seam starved one whole fact kind at trivially small store sizes -- buildHintFormat splits results into an aggressive set (preferences and corrections, cap 8) and a precision set (everything else, cap 4), but it called retrieve() without a limit, so it received DEFAULT_RECALL_LIMIT (20) results and split those. That limit is a post-ranking slice applied before the kind-split, so whenever the top 20 shared one kind the other cap was starved to zero. It could never have bounded the output anyway -- the caps total 12, already under 20 -- so its only effect was to distort composition.

    Score ties break on captured_at descending, which makes the triggering shape an ordinary one rather than an extreme one: older decisions, newer preferences accumulating on top. Reproduced against the built bundle on a 33-fact store (8 decisions recorded first, then 25 preferences): the pre-fix binary emitted 8 preference lines and zero decision lines, the post-fix binary emits 8 and 4, from the same database and the same command. Every decision the user had recorded was invisible to their agent, in a payload carrying the normal header, grammar and footer -- indistinguishable from "this project has no decisions."

    This is the third path in the class the entry above describes, and the one that reaches a real user first: it needs no timing pressure and no unusual scale, only more recent facts of one kind than the recall limit. HINT_FORMAT_RECALL_LIMIT now makes the seam's own caps the sole bound on what reaches the wire. Found by the new scale-invariant test rather than by hand -- the assertion that 500 seeded facts emit exactly 8 + 4 failed at 8 + 0.

  • Nothing exercised the seam at a store size where its caps bind -- every existing test seeded a handful of facts, so the emission caps, the kind-split and the budget path were only ever tested where they could not interact. tests/unit/integration-seam.test.ts now seeds 500 facts and asserts the invariant directly: the fact-line count is the full cap set or zero, never a third size, with the wire-shape assertion that a complete set carries the footer and an empty one carries nothing. The third of those tests deliberately runs against the default budget so it exercises whichever path the runner takes, which is what makes it an invariant guard rather than a latency assertion -- recall is linear in store size, and a wall-clock assertion on a shared runner is the flake this file already carries a wrapper to prevent.

  • The budget-pinning convention was a convention, and one file had already missed it -- tests/unit/integration-seam.test.ts gained a wrapper defaulting the retrieval soft budget in 0.3.0; tests/integration-seam.test.ts was never given one. That stayed survivable only while budget exhaustion shrank the result, since a content assertion could still pass against the smaller set. Once exhaustion emptied the result instead, the same latent flake turned two Windows CI jobs red. tests/guards/seam-budget.test.ts now fails if any test file imports buildHintFormat without defining a pin, and carries a second assertion that at least one file is being watched, so a rename cannot quietly reduce it to a test that asserts an empty list is empty.

  • 26 tests asserted an error class where the class could not identify the failure -- expect(...).toThrow(SomeError) passes for any instance of that class from any code path, so wherever one class covers several distinct failures the assertion silently weakens into "something went wrong." WiringConflictError carries 14 distinct messages and CaptureValidationError 19; a test named for one guard was free to pass on any of the others. This was not hypothetical: 0.3.0's own oversized-import test, named "throws JsonImportError before attempting to parse," was passing on the parse -- 50MB of filler is not valid JSON either, and raising the size limit tenfold left the test green, meaning the guard it existed for had never been covered.

    Every one of the 26 now pins the message alongside the class. Auditing them found no further live defects -- each was throwing what its name claimed -- but four sites in tests/capture.test.ts and tests/unit/capture.test.ts were missed by the manual sweep and caught only by the new guard, which is the argument for having one. tests/guards/error-assertions.test.ts fails when a bare class assertion names a class src/ can raise with more than one message. Its ambiguity model counts construction sites rather than throw sites, because readFileWithErrorMapping (src/fileUtils.ts) constructs the caller's class across four message branches -- which is how JsonImportError reaches nine possible messages from five direct throws, and why MarkdownImportError, never thrown directly at all, would otherwise have counted as unambiguous. Classes defined inside a test file are exempt: tests/unit/fileUtils.test.ts passes TestError in to prove the mapper returns the class it was given, so there the class is the assertion.

    Verified in both directions. Stripping one message assertion makes the guard fail naming the file, class and message count -- while the 85-test wiring suite it came from stays entirely green, which is precisely the regression the guard exists to catch and the suite cannot see.

  • mem recall <query> returned the whole store for a query that matched nothing -- byte-identical to mem recall with no query at all. A query is a ranking input, not a filter: BM25 orders the candidate set and never removes from it, so the results.length === 0 branch that prints no matching facts could only ever fire when a --kind/--scope filter excluded everything, never when the query itself matched nothing. On a three-fact store, mem recall xyzzyplughquux returned all three facts, ordered by recency, presented exactly as a hit. The no matching facts outcome src/cli.ts documents was unreachable by the query path.

    Recall now says so and still shows the facts: the fix adds the missing signal rather than emptying the result. The condition is that every result scored 0, which is the documented meaning of an empty query ("all candidates tie at score 0", src/retrieval.ts) and also covers a term so common it appears in every fact -- zero discriminating power, so "did not narrow these results" is true there too, which is why the wording claims that rather than claiming the term is absent. The note is human-path only: --hint-format returns before it, so the closed TGMEM/2 grammar is untouched, and a regression test asserts no such line reaches the wire.

  • mem list --kind decision reported "no facts stored" on a store that was not empty -- the message is a claim about the whole store, and a filter excluding everything is a different fact about the world. A user filtering a populated store was told it held nothing. It now distinguishes the two, and still says "no facts stored" when the store really is empty.

  • exitOverride() covered the root command only, so all 15 subcommands exited mid-flush -- run() sets it so Commander cannot "call process.exit() mid-flush", per the comment that has always sat above the line. It is not inherited: a Command applies it to itself alone. Every subcommand therefore kept Commander's default behaviour and called process.exit() directly on any parse error or --help of its own -- the exact mid-flush exit the line exists to prevent, and the same class as the EPIPE truncation fixed in 0.3.0.

    It also made the handler's own commander.-prefixed branch -- which names "invalid option" in its comment -- unreachable for every subcommand. A subcommand parse failure surfaced as a plain Error with no code, so it was classified as an internal bug and exited 2 where the contract specifies 1 for a usage error. Production masked this completely: Commander's own process.exit(1) produced the correct code before mem's handler ran, so only the in-process path could reveal it, and mem list --help was returning a failure code there. Fixed by extending exitOverride() to every registered subcommand; verified that root and subcommand help and version still exit 0, on both the in-process path and the built bundle.

  • Nothing checked that a command which did nothing said so -- src/cli.ts makes exit 0 normative for "nothing found" outcomes, which is a deliberate and correct position, but it means the exit code carries no signal for an empty result and the whole burden falls on stdout. Nothing tested that stdout actually carried it. tests/no-signal-sweep.test.ts now walks every registered command in a state where nothing succeeded and asserts the signal that makes that state legible -- the exit code where the outcome is a user error, a specific stdout line where the contract makes it a success. All three defects above were found by writing it, and 8 of its 24 tests were confirmed failing before the fixes.

  • src/anchors.ts had the suite's weakest coverage on its most adversarial code -- 85.31% of statements and 83.38% of branches, the lowest of any src/ module, on the one file whose whole job is to parse hostile input: a binary .git/index written by another program, glob patterns from a stored fact, and paths that may point anywhere. Every uncovered branch was a malformed-input or refusal path -- a truncated index header, an entry whose name length runs past the buffer, a gitdir: pointer that escapes the root, a symlinked .git -- so the untested set was precisely the set that decides whether a corrupt or adversarial repository yields unverified or a wrong verdict. Anchors are the input to freshness, and a wrong affirmed presents a stale fact as ground truth.

    tests/anchors-branches.test.ts covers them: 59 tests taking src/anchors.ts to 95.10% statements and 96.00% branches, with every function hit. It builds .git/index files byte by byte -- header magic, version, entry count, the 62-byte metadata block, the v3 extended-flag bit, the 8-byte padding -- so each malformed shape is a deliberate single deviation from a valid file rather than random noise, and asserts unverified for all of them. The budget-exhaustion tests drive a deterministic clock and additionally assert that a later unbudgeted call returns a real verdict, proving the bailout was not memoized into the verdict cache.

    No production defect was found. Every path already behaved correctly; this closes a test gap, not a bug. The remaining uncovered branches are individually accounted for rather than left unexplained: five symlink-refusal tests are skipIf(isWindows) and run on the Linux CI jobs, one branch is a per-platform ternary that only the CI matrix can cover both sides of, two are the 20,000-entry glob cap, and the rest are narrowing guards a preceding regex or arity check already makes unreachable. Whole-suite coverage rises to 94.23% statements / 87.89% branches / 98.31% functions / 94.10% lines.

  • mem show --json told every consumer a fact had no recorded sources, when mem records none for any fact -- the envelope carries a sources array, and because nothing in the capture path ever writes a source row, that array is [] for every fact in every install. [] is the same value a fact with genuinely no sources would produce, so a consumer had no way to tell "this fact has none" from "mem has none, ever" -- a well-formed payload indistinguishable from a truthful one, the same class of defect as a seam that withholds facts while its output claims completeness. mem show --help made it worse by advertising the field as a capability: "Includes freshness and sources, unlike mem export/mem list --json", naming an always-empty array alongside a verdict that is always real. The honest description existed, but only in AGENTS.md and CLAUDE.md -- the two agent-facing files. Both user-facing surfaces, the README table and the CLI's own help text, claimed the capability without the caveat.

    Both now say the array is reserved and always empty, and say what [] therefore means. The human mem show needed no change: it omits the sources: block entirely when there are none, so it never made the claim.

  • The sources table stays, and that is now a recorded decision rather than an unexplained gap -- the table, its storage API and its gc pruning are fully implemented and tested with no caller, which invites deletion as dead code. Deleting it is a schema migration plus a break of three exports in src/index.ts, a one-way door bought for no correctness gain; feeding it is a new feature. Neither is worth doing, so it is kept as an explicitly unfed seam and AGENTS.md records why.

    tests/guards/unfed-sources.test.ts holds that state to the code in both directions: while no src/ file calls insertSource, every surface mentioning sources must carry the "always empty" disclosure; the moment one does, the guard fails and names the disclosures that have just become false. Both directions were confirmed failing before being relied on -- stripping the README caveat fails the first test by name, adding a single insertSource call to the capture path fails the second with its file and line.