Repository navigation
v0.3.1
Fixed
-
The agent-facing seam withheld facts while its output claimed to be complete -- when
buildHintFormatexceeded its 150 ms soft budget it dropped its emission caps from 8/4 to 2/1 and returned the smaller set. Measured against a real store: 500 facts returned 12 lines, 2,000 returned 12, 10,000 returned 8, and 40,000 returned 2. The payload at 40,000 was byte-shape-identical to a healthy one -- sameTGMEM/2header, same line grammar, same footer -- so a consumer had no way to tell 2 facts from all of them.HintFormatResult.truncatedwas computed and returned, butsrc/cli.tsnever read it, and the only other signal was alogWarningon stderr, which the documented consumer contract does not read: README describes exactly two consumer failure modes, a missing binary and a caller timeout, and this is neither. Truncation fires when mem overruns its own budget and still answers inside the caller's window, so the caller's fail-open path never engages. For a seam whose stated guarantee is self-caveating output, that is the one failure it is not allowed to have.There is no way to say "this is partial" in TGMEM/2. The grammar is closed: a conforming consumer drops an off-grammar line rather than guessing at it,
footer-textis pinned to one exact string, and any grammar change bumps the version -- at which point consumers that have not upgraded fail open to no hints at all. So annotating a reduced response is not available, and a reduced response cannot be distinguished from a complete one. The budget-exhausted path now returns an empty hint set instead, which is already this module's shape for "I could not deliver" (buildHintFormat's internal-failure catch returns the same thing) and which the consumer's fail-open path already handles.TRUNCATED_AGGRESSIVE_CAPandTRUNCATED_PRECISION_CAPare gone; no wire change and no version bump were needed.Verified end to end against the built bundle at 40,001 facts over six runs: output is now strictly binary -- 10 lines (header, 8 fact-lines, footer) or 1 line (bare
TGMEM/2), never a subset. Both regression tests were confirmed failing before the fix, emitting 4 and 2 lines respectively. The in-process test assertslinesis exactly empty rather than merely shorter than a healthy response, because ashorter-thanassertion would pass again the moment a reduced cap was reintroduced, which is the defect; the wire-level test asserts no fact-lines and no footer reach stdout, since TGMEM/2 emits a footer only alongside at least one fact-line and a lone footer would itself be off-grammar. -
The seam starved one whole fact kind at trivially small store sizes --
buildHintFormatsplits results into an aggressive set (preferences and corrections, cap 8) and a precision set (everything else, cap 4), but it calledretrieve()without a limit, so it receivedDEFAULT_RECALL_LIMIT(20) results and split those. That limit is a post-ranking slice applied before the kind-split, so whenever the top 20 shared one kind the other cap was starved to zero. It could never have bounded the output anyway -- the caps total 12, already under 20 -- so its only effect was to distort composition.Score ties break on
captured_atdescending, which makes the triggering shape an ordinary one rather than an extreme one: older decisions, newer preferences accumulating on top. Reproduced against the built bundle on a 33-fact store (8 decisions recorded first, then 25 preferences): the pre-fix binary emitted 8 preference lines and zero decision lines, the post-fix binary emits 8 and 4, from the same database and the same command. Every decision the user had recorded was invisible to their agent, in a payload carrying the normal header, grammar and footer -- indistinguishable from "this project has no decisions."This is the third path in the class the entry above describes, and the one that reaches a real user first: it needs no timing pressure and no unusual scale, only more recent facts of one kind than the recall limit.
HINT_FORMAT_RECALL_LIMITnow makes the seam's own caps the sole bound on what reaches the wire. Found by the new scale-invariant test rather than by hand -- the assertion that 500 seeded facts emit exactly 8 + 4 failed at 8 + 0. -
Nothing exercised the seam at a store size where its caps bind -- every existing test seeded a handful of facts, so the emission caps, the kind-split and the budget path were only ever tested where they could not interact.
tests/unit/integration-seam.test.tsnow seeds 500 facts and asserts the invariant directly: the fact-line count is the full cap set or zero, never a third size, with the wire-shape assertion that a complete set carries the footer and an empty one carries nothing. The third of those tests deliberately runs against the default budget so it exercises whichever path the runner takes, which is what makes it an invariant guard rather than a latency assertion -- recall is linear in store size, and a wall-clock assertion on a shared runner is the flake this file already carries a wrapper to prevent. -
The budget-pinning convention was a convention, and one file had already missed it --
tests/unit/integration-seam.test.tsgained a wrapper defaulting the retrieval soft budget in 0.3.0;tests/integration-seam.test.tswas never given one. That stayed survivable only while budget exhaustion shrank the result, since a content assertion could still pass against the smaller set. Once exhaustion emptied the result instead, the same latent flake turned two Windows CI jobs red.tests/guards/seam-budget.test.tsnow fails if any test file importsbuildHintFormatwithout defining a pin, and carries a second assertion that at least one file is being watched, so a rename cannot quietly reduce it to a test that asserts an empty list is empty. -
26 tests asserted an error class where the class could not identify the failure --
expect(...).toThrow(SomeError)passes for any instance of that class from any code path, so wherever one class covers several distinct failures the assertion silently weakens into "something went wrong."WiringConflictErrorcarries 14 distinct messages andCaptureValidationError19; a test named for one guard was free to pass on any of the others. This was not hypothetical: 0.3.0's own oversized-import test, named "throws JsonImportError before attempting to parse," was passing on the parse -- 50MB of filler is not valid JSON either, and raising the size limit tenfold left the test green, meaning the guard it existed for had never been covered.Every one of the 26 now pins the message alongside the class. Auditing them found no further live defects -- each was throwing what its name claimed -- but four sites in
tests/capture.test.tsandtests/unit/capture.test.tswere missed by the manual sweep and caught only by the new guard, which is the argument for having one.tests/guards/error-assertions.test.tsfails when a bare class assertion names a classsrc/can raise with more than one message. Its ambiguity model counts construction sites rather thanthrowsites, becausereadFileWithErrorMapping(src/fileUtils.ts) constructs the caller's class across four message branches -- which is howJsonImportErrorreaches nine possible messages from five direct throws, and whyMarkdownImportError, never thrown directly at all, would otherwise have counted as unambiguous. Classes defined inside a test file are exempt:tests/unit/fileUtils.test.tspassesTestErrorin to prove the mapper returns the class it was given, so there the class is the assertion.Verified in both directions. Stripping one message assertion makes the guard fail naming the file, class and message count -- while the 85-test
wiringsuite it came from stays entirely green, which is precisely the regression the guard exists to catch and the suite cannot see. -
mem recall <query>returned the whole store for a query that matched nothing -- byte-identical tomem recallwith no query at all. A query is a ranking input, not a filter: BM25 orders the candidate set and never removes from it, so theresults.length === 0branch that printsno matching factscould only ever fire when a--kind/--scopefilter excluded everything, never when the query itself matched nothing. On a three-fact store,mem recall xyzzyplughquuxreturned all three facts, ordered by recency, presented exactly as a hit. Theno matching factsoutcomesrc/cli.tsdocuments was unreachable by the query path.Recall now says so and still shows the facts: the fix adds the missing signal rather than emptying the result. The condition is that every result scored 0, which is the documented meaning of an empty query ("all candidates tie at score 0",
src/retrieval.ts) and also covers a term so common it appears in every fact -- zero discriminating power, so "did not narrow these results" is true there too, which is why the wording claims that rather than claiming the term is absent. The note is human-path only:--hint-formatreturns before it, so the closed TGMEM/2 grammar is untouched, and a regression test asserts no such line reaches the wire. -
mem list --kind decisionreported "no facts stored" on a store that was not empty -- the message is a claim about the whole store, and a filter excluding everything is a different fact about the world. A user filtering a populated store was told it held nothing. It now distinguishes the two, and still says "no facts stored" when the store really is empty. -
exitOverride()covered the root command only, so all 15 subcommands exited mid-flush --run()sets it so Commander cannot "callprocess.exit()mid-flush", per the comment that has always sat above the line. It is not inherited: aCommandapplies it to itself alone. Every subcommand therefore kept Commander's default behaviour and calledprocess.exit()directly on any parse error or--helpof its own -- the exact mid-flush exit the line exists to prevent, and the same class as the EPIPE truncation fixed in 0.3.0.It also made the handler's own
commander.-prefixed branch -- which names "invalid option" in its comment -- unreachable for every subcommand. A subcommand parse failure surfaced as a plainErrorwith nocode, so it was classified as an internal bug and exited 2 where the contract specifies 1 for a usage error. Production masked this completely: Commander's ownprocess.exit(1)produced the correct code before mem's handler ran, so only the in-process path could reveal it, andmem list --helpwas returning a failure code there. Fixed by extendingexitOverride()to every registered subcommand; verified that root and subcommand help and version still exit 0, on both the in-process path and the built bundle. -
Nothing checked that a command which did nothing said so --
src/cli.tsmakes exit 0 normative for "nothing found" outcomes, which is a deliberate and correct position, but it means the exit code carries no signal for an empty result and the whole burden falls on stdout. Nothing tested that stdout actually carried it.tests/no-signal-sweep.test.tsnow walks every registered command in a state where nothing succeeded and asserts the signal that makes that state legible -- the exit code where the outcome is a user error, a specific stdout line where the contract makes it a success. All three defects above were found by writing it, and 8 of its 24 tests were confirmed failing before the fixes. -
src/anchors.tshad the suite's weakest coverage on its most adversarial code -- 85.31% of statements and 83.38% of branches, the lowest of anysrc/module, on the one file whose whole job is to parse hostile input: a binary.git/indexwritten by another program, glob patterns from a stored fact, and paths that may point anywhere. Every uncovered branch was a malformed-input or refusal path -- a truncated index header, an entry whose name length runs past the buffer, agitdir:pointer that escapes the root, a symlinked.git-- so the untested set was precisely the set that decides whether a corrupt or adversarial repository yieldsunverifiedor a wrong verdict. Anchors are the input to freshness, and a wrongaffirmedpresents a stale fact as ground truth.tests/anchors-branches.test.tscovers them: 59 tests takingsrc/anchors.tsto 95.10% statements and 96.00% branches, with every function hit. It builds.git/indexfiles byte by byte -- header magic, version, entry count, the 62-byte metadata block, the v3 extended-flag bit, the 8-byte padding -- so each malformed shape is a deliberate single deviation from a valid file rather than random noise, and assertsunverifiedfor all of them. The budget-exhaustion tests drive a deterministic clock and additionally assert that a later unbudgeted call returns a real verdict, proving the bailout was not memoized into the verdict cache.No production defect was found. Every path already behaved correctly; this closes a test gap, not a bug. The remaining uncovered branches are individually accounted for rather than left unexplained: five symlink-refusal tests are
skipIf(isWindows)and run on the Linux CI jobs, one branch is a per-platform ternary that only the CI matrix can cover both sides of, two are the 20,000-entry glob cap, and the rest are narrowing guards a preceding regex or arity check already makes unreachable. Whole-suite coverage rises to 94.23% statements / 87.89% branches / 98.31% functions / 94.10% lines. -
mem show --jsontold every consumer a fact had no recorded sources, when mem records none for any fact -- the envelope carries asourcesarray, and because nothing in the capture path ever writes a source row, that array is[]for every fact in every install.[]is the same value a fact with genuinely no sources would produce, so a consumer had no way to tell "this fact has none" from "mem has none, ever" -- a well-formed payload indistinguishable from a truthful one, the same class of defect as a seam that withholds facts while its output claims completeness.mem show --helpmade it worse by advertising the field as a capability: "Includes freshness and sources, unlike mem export/mem list --json", naming an always-empty array alongside a verdict that is always real. The honest description existed, but only inAGENTS.mdandCLAUDE.md-- the two agent-facing files. Both user-facing surfaces, the README table and the CLI's own help text, claimed the capability without the caveat.Both now say the array is reserved and always empty, and say what
[]therefore means. The humanmem showneeded no change: it omits thesources:block entirely when there are none, so it never made the claim. -
The
sourcestable stays, and that is now a recorded decision rather than an unexplained gap -- the table, its storage API and its gc pruning are fully implemented and tested with no caller, which invites deletion as dead code. Deleting it is a schema migration plus a break of three exports insrc/index.ts, a one-way door bought for no correctness gain; feeding it is a new feature. Neither is worth doing, so it is kept as an explicitly unfed seam andAGENTS.mdrecords why.tests/guards/unfed-sources.test.tsholds that state to the code in both directions: while nosrc/file callsinsertSource, every surface mentioningsourcesmust carry the "always empty" disclosure; the moment one does, the guard fails and names the disclosures that have just become false. Both directions were confirmed failing before being relied on -- stripping the README caveat fails the first test by name, adding a singleinsertSourcecall to the capture path fails the second with its file and line.