Skip to content

Memtrace v1.1.9

Choose a tag to compare

@Alex793x Alex793x released this 03 Sep 03:17
· 51 commits to main since this release

Memtrace v1.1.9

Memtrace governs its own CPU

New

  • Memtrace now governs its own CPU. Background work — daemon rebuilds, the file watcher, embedding, cross-repo convergence — draws from one CPU budget (half the performance cores by default) that yields to your own builds and editors, runs at a quarter rate on battery and pauses below 15%, and lifts to full speed once you have been away from the keyboard for five minutes. An interactive memtrace index keeps every performance core but still backs off under heat, on battery, or under critical memory pressure. MEMTRACE_CPU_GOVERNOR=off restores the previous scheduling; MEMTRACE_CPU_BUDGET_CORES sets the allowance and MEMTRACE_BACKGROUND_ON_BATTERY=off|full changes the battery policy.
  • An idle-burn watchdog ends the class of overheating where Memtrace burned CPU for hours on a machine where nothing changed. When the daemon and its database sidecar spend CPU for ten minutes without parsing, embedding, converging, or serving anything, every background loop is parked and the budget drops to a trickle until real work arrives — a saved file or an index request lifts it immediately. memtrace status, /api/health, and the mem_diag tool now show where every CPU-second went (parse, embed, convergence, watcher, serving), the current budget and why it was chosen, and any parked loops.
  • New features land on Nightly first: npm install -g memtrace@nightly is the same memtrace command, and it is not always stable. They still ship on the default install. memtrace install stays on the channel you already have.
  • A real int8 build of the default code model is available as an opt-in (MEMTRACE_EMBED_QUANT=q8): a 162 MB first-run download instead of 641 MB and the same recall on natural-language queries against the vector index alone. In a paired benchmark of name-style find_code queries the int8 graph combined with the new length-sorted batching lost 4 points of top-10 accuracy (the quantised graph calibrates per batch, so documents batched together drift away from queries embedded alone), while the int8 graph on its own lost half a point at 45 % less embedding CPU. So every host keeps the full-precision graph by default, batching is switched off automatically under q8, and existing caches stay valid; switching to q8 re-embeds once. (The auto tier's previous "int8" label had always loaded the full-precision graph; the label is unchanged.)
  • Code the machine has already embedded is never embedded again: a second worktree, a clone under another name, a fork, or a vendored copy of the same symbols reuses the vectors from a machine-wide cache keyed by symbol content (MEMTRACE_EMBED_CACHE_GLOBAL_CAP_MB, default 512 MB). MEMTRACE_EMBED_CACHE=off disables caching for benchmarks.

Improved

  • Bulk indexing threads, the embedding worker, and the database sidecar now run below normal priority — on Windows 11 with EcoQoS, which moves them onto efficiency cores — so Memtrace stays out of the way of whatever you are doing. MEMTRACE_SIDECAR_PRIORITY=normal keeps the sidecar at the default priority.
  • Embedding, reranking, and SPLADE inference threads no longer spin-wait between operators, which burned one core per thread whenever the model pool was idle; MEMTRACE_ORT_SPIN=1 restores spinning. The SPLADE session also no longer takes one thread per core.
  • Hybrid Intel and AMD processors are now recognised on Windows and Linux. An 8P+12E Core Ultra previously counted as 20 performance cores, so both the interactive and the background thread pools sized themselves to every core on the chip.
  • Embedding batches are sorted by length and cut by a token budget, so short symbols no longer pay for the longest symbol in their batch (every batch used to be padded to its longest text). The vectors are identical to the unsorted ones (per-document cosine 1.00000 in our check) at 28 to 35 % less embedding CPU per symbol in our measurements; memtrace status and /api/health report the share of inference spent on real tokens. MEMTRACE_EMBED_BUCKETING=off restores the old batching.
  • The vector index writes far less to disk: adding a symbol no longer rewrites the full record (vector included) of every neighbour it links to, only their neighbour lists, and the 1-bit codec (MEMTRACE_HNSW_CODEC=rabitq) now scores candidates by popcount instead of decoding each one back to floats. The first start after this upgrade rebuilds the vector index once from the stored vectors.
  • A new embed-text recipe is available for operators who want the last bit of indexing CPU: MEMTRACE_EMBED_RECIPE=headtail:900,400 embeds the head and the tail of long symbol bodies instead of the first 1500 characters — in our measurements the same recall at 13 % less embedding CPU. It changes the cache namespace, so switching re-embeds once; the default recipe is unchanged.
  • memtrace gpu status reports the NPU: whether one is present, whether the OpenVINO runtime and provider are installed, and what is still missing before MEMTRACE_ENABLE_OPENVINO=1 can use it. MEMTRACE_OPENVINO_DEVICE selects the device (default AUTO:NPU,CPU) and the provider now takes precedence over DirectML when enabled.
  • On Windows, a failed ONNX Runtime load now names the usual cause — the Microsoft Visual C++ 2015-2022 runtime is missing — with the download, instead of pointing at PATH.
  • memtrace doctor on Windows now checks for the Microsoft Visual C++ 2015-2022 runtime that Memtrace's embedding runtime needs and prints the download link when it is missing, so a fresh machine learns this before the first memtrace start fails to load ONNX Runtime.
  • list_watched_paths, embed_diag and mem_diag now say which process answered (_meta.scope: the daemon, or an MCP session's own stdio process, with the daemon's pid), so an empty watch list or idle embed state seen from an editor session is no longer mistaken for the daemon doing nothing.
  • MCP tool results say when the graph is still being written. While indexing or cross-repo linking is running, every result carries _meta.graph_converging: true with the pass state, so an agent can tell a partial impact or caller answer from a settled one and retry instead of trusting it. Thanks to Magalz for the report.

Fixed

  • A working-tree change whose exact delta could not be computed re-parsed the repository every five seconds forever, and every attempt re-armed cross-repo linking. The retry now backs off from five seconds to five minutes and resets as soon as a delta applies.
  • The observe-mode measurement drain no longer runs real searches every 20 seconds against an empty spool; it backs off and parks after five empty passes, resuming on its own when entries appear.
  • A vector index holding a single symbol was never written to disk until a second symbol arrived, and compacting a 1-bit index silently rebuilt it as full floats.
  • Semantic search kept working on large stores: once a local store passed 512 MB, background maintenance of the vector index was switched off to save CPU, so every symbol saved after that was findable by name but not by meaning. The vector index now keeps updating on stores of any size; only the structural index rebuild is skipped on large stores, and memtrace start says so.
  • On Windows the file watcher reported itself armed but never recorded a change: the watched folder was stored in one path spelling and change events arrived in another, so every save was silently dropped. Saves and pulls now reach the index, dropped events are counted, and a batch that lands during a merge or rebase is deferred until git is done instead of being thrown away.
  • memtrace govern and the dashboard's Governance panel no longer fail with "cortex.sock I/O failed: The pipe is being closed" on repositories with many rules or ADRs (Windows). Classification is now sent in bounded batches instead of one request carrying every document's text.
  • memtrace govern docs/architecture now finds ADRs under numbered or underscored directories such as 09_architecture_decisions; previously the sweep reported 0 governance docs for them.
  • The installer registers the Memtrace MCP server where Claude Code actually reads it (~/.claude.json) when the claude CLI is unavailable, and memtrace doctor checks that file; before, skills appeared in a session while every Memtrace tool was missing and doctor still reported the registration healthy.
  • memtrace start picks the next free port when 3030 turned out busy between the check and the bind instead of aborting.
  • memtrace govern and the Governance panel no longer say "MemCortex is not running" when MemCortex is running but did not accept a connection in time (starting up or busy); the message now says what happened and suggests retrying.
  • Restarting Memtrace on a large store no longer burns CPU for an hour with nothing to do. After an unclean stop (a crash, a kill, a machine restart) the database used to flush every mapped page of the whole store up to three times and rewrite its entire record index, and when the vector-index log had grown past 512 MB it replayed that whole log in the background on every start, forever, because nothing ever shrank it. The checkpoint now costs only what changed since the last one, and the vector-index log is rotated down to one record per live symbol after such a recovery and whenever it outgrows the live index, so the next start recovers it in seconds. Measured on a copied store restarted after a kill: checkpoint 0.12 s, vector index recovered from the rotated log with no background rebuild. Existing stores get the rotation on their first start after the upgrade.
  • Restarting Memtrace no longer re-embeds every repository. Three minutes after each start, a background clean-up of duplicate record versions marked every repository's symbols as changed even when it had nothing to close, which invalidated the embedding completion record (or withheld it while embedding was still running), so the next start re-ran the whole embedding pass — on large stores, sustained CPU after every restart. The clean-up now scans first and records a change only when it actually closes something.
  • memtrace mcp answers initialize as soon as it is attached to the workspace store, before it hydrates repositories, starts Cortex, or validates the licence session; those now run in the background and the first tool call waits for them (up to a minute). On a busy or very large workspace the server used to sit silently after MemDB ready until every MCP client timed out. While it starts, stderr now names the phase it is in (still starting (10s): validating the licence session), and a start that cannot reach the session within two minutes exits with that phase instead of lingering — so periodic probes no longer accumulate as hung processes. MEMTRACE_MCP_STARTUP_TIMEOUT_SECS raises the limit. Thanks to BadMrPotatoHead for the report.
  • A startup embedding pass no longer re-embeds symbols that already have a vector. When a restart found the embedding completion record missing, the daemon re-embedded the entire repository from cache and rewrote every vector in the store — hours of sustained CPU on a large workspace, on every start, and it never converged when the pass was interrupted. The pass now skips symbols whose vector already exists, records those as complete, and so finishes in proportion to what is actually missing. memtrace index now reports the symbols that already had vectors in its summary, and MEMTRACE_FORCE_EMBED=1 re-embeds everything, for a model change or a suspect store. Thanks to BadMrPotatoHead for the report.
  • On Windows, memtrace doctor --fix, memtrace reset, and memtrace stop now recognise that a workspace owner has exited. The liveness check behind all three always answered "alive", so a killed owner's daemon.pid and daemon-state.json were never swept, doctor reported "heartbeat is alive but health is not responding", and reset refused to proceed until the files were moved by hand. A pid is now checked against its real exit state, including the case where a parent shell still holds a handle to the dead process. Thanks to Magalz for the report.
  • Related symbols in get_symbol_context (contains, community_siblings, callers) now carry annotations and annotation_args like the primary symbol; the related-symbol projection had dropped both fields, which rendered them as null for symbols whose direct payload had them. Thanks to Magalz for the report.
  • get_repository_stats.nodes_by_kind now includes every kind the indexer stamps — Constant, Field, Variable, TypeAlias, SQL, Terraform, and CI kinds among them — instead of a fixed subset, so the rollup reconciles with what find_symbol returns. Thanks to Magalz for the report.
  • Attaching memtrace mcp to a long-running owner no longer fails after the machine's clock is adjusted. The owner's runtime record was verified by its wall-clock start time, which shifts under WSL2 after a clock step and made every later attach refuse a healthy daemon ("has no runtime record for store"). On Linux the record now carries the kernel's clock-tick start time and boot id and is verified by those; other platforms keep the start-time check. Thanks to Magalz for the report.
  • Agent guidance from the Memtrace MCP server now reaches the agent intact. The routing instructions had grown past the size limit coding agents apply to a server's instructions, so the tail was silently cut off — including the note telling agents to pass an explicit repo_id in an ambiguous multi-repo workspace, which is exactly where scoping mistakes happen.
  • Memtrace's MCP tools are now offered to the agent from the first turn. Coding agents defer tool definitions until something asks for them, so Memtrace's tools arrived as bare names while the built-in search tools arrived complete — biasing the agent toward grep before it could see what Memtrace offered. Run memtrace install to pick this up.
  • Operator surfaces now answer for the daemon that actually does the work. An MCP child spawned without workspace context (an IDE launching at the filesystem root) used to answer list_jobs, list_watched_paths, mem_diag, and embed_diag from its own empty process — zero watches while the daemon held 44, an 18 MB memory report while the engine held gigabytes — with no hint anything was wrong. These surfaces now find the workspace daemon through the same runtime discovery repository listing uses and answer from it, naming the answering process in _meta.answered_by; jobs and watches merge both processes' entries, a daemon job id can be polled from the child, and when no daemon is reachable the payload says so (_meta.daemon_unreachable with the exact reason and remedy — including "restart the daemon" when it predates these routes) instead of presenting silent empty lists as the truth.

Install

npm install -g memtrace