Skip to content

v3.1.0 — smem index/reindex batch their writes and stop duplicating on re-run

Choose a tag to compare

@acidkill acidkill released this 03 Aug 13:00
· 109 commits to main since this release
4b57298

Added

  • smem index --force — wipes the existing code index and rebuilds it from scratch, ignoring change tracking. The explicit escape hatch for the new incremental-skip behaviour below. (#146)

Fixed

  • smem index duplicated every neuron/synapse/fiber on every re-run. CodebaseEncoder never hashed or compared file content, so re-indexing an unchanged tree recreated it wholesale with fresh IDs. It now tracks each file's mtime and a content simhash on the file's own neuron: unchanged files are skipped entirely, a touched-but-content-identical file is skipped via the simhash comparison, and a genuinely changed file has its previous neurons/synapses/fiber removed before being rebuilt. (#146)
  • [embedding] endpoint in config.toml was silently dead. _create_provider never called EmbeddingSettings.resolved_endpoint(), so the openai and ollama providers only ever picked up an endpoint from the SURREAL_MEMORY_EMBEDDING_ENDPOINT environment variable — the config field had no effect on the canonical embedding path (smem_remember, smem reindex, the doc trainer). Both providers now receive the resolved endpoint, and the provider cache key includes it so editing the endpoint returns a freshly built provider instead of one pointed at the old address. (#146)
  • smem doctor's "full setup guide" link and the generated IDE-rules file both pointed at a stale fork username (nhadaututtheky instead of acidkill) — leftover from a fork/rename that was never fully updated. (#155)
  • The S310 lint suppression in mcp/version_check_handler.py broke under ruff < 0.16 — the diagnostic's anchor line moved between ruff versions; naming both S310 and RUF100 on the same directive satisfies both. (#147, contributed by @RobertSigmundsson)
  • A unit test failed on any machine whose TMPDIR is under the home directory (test_relocated_global_config_dir_is_not_a_project_root) — the patched home fixture was a sibling of the patched cwd rather than an ancestor, so detect_project_root()'s upward walk left the fixture entirely and picked up a real marker from the actual home directory. (#148, contributed by @RobertSigmundsson)
  • pre-commit run --all-files failed 7 of 12 hooks on a clean checkout, and a plain git commit failed 4 of them — nobody could follow CONTRIBUTING.md's own setup instructions. The ruff hook was pinned below the version that could parse pyproject.toml's own ignore list, the mypy hook ran without the project's runtime dependencies (producing 43 false positives), check-yaml couldn't load mkdocs.yml's mkdocs-material tags, and the whitespace hooks were rewriting bytes inside the committed dashboard bundle's template literals. Hook versions raised, scopes corrected, and four source JSON/config files given their missing final newline. (#149, contributed by @RobertSigmundsson)

Performance

  • smem index's directory walk stopped materializing the whole tree before filtering it. It used to sorted(directory.rglob("*")) over the entire tree — including excluded directories like .git, .venv, and nested .claude/worktrees checkouts — and only rejected them file-by-file afterward, with a resolve() syscall on every single path before the extension/exclude filters ran. It now walks with os.walk, pruning excluded directory names in place so they're never descended into, and only resolves paths that already passed the extension filter. (#146)
  • Codebase indexing batches its database writes. Every file used to cost 3 sequential round-trips per neuron (entity + activation state + change-log row), 2 per synapse, and 2 per fiber. CodebaseEncoder now builds every file's neurons/synapses/fiber in memory across the whole directory and flushes them in a handful of multi-statement writes at the end (new add_neurons_batch/add_fibers_batch on NeuralStorage, with a sequential fallback for backends that don't override them). add_synapses_batch's own change-log write was also fixed to log in one bulk insert instead of looping per synapse — a "batch" writer that still cost N round-trips for its own bookkeeping. (#146)
  • smem reindex writes embedding vectors in one round-trip per batch instead of one per neuron, via the existing update_neuron_embeddings batch writer (now promoted to a NeuralStorage interface method with a sequential fallback, replacing an ad hoc getattr check). (#146)

Changed

  • Embeddings for code symbols are computed from more than just the symbol's name. A code neuron's content doubles as its dedup/identity key (Klasa.metoda), which left almost nothing for the embedding provider to work with. Neuron.embedding_text() composes content with the symbol's signature/docstring from its metadata when present, and both smem reindex and inline post-encode embedding now use it — content itself, and dedup, are unaffected. (#146)
  • Bandit's security-lint suppression list was brought back in line with what ruff already ignores project-wide, plus documented, targeted exclusions for SurrealQL query construction (identifiers pass a strict charset validator; values go through $-bindings) and deliberate fail-soft try/except paths. Nine individual findings — enum labels naming secret types rather than containing one, a documented default-credential fallback, an escape-only XML import, and similar — get a per-line justification instead of a blanket skip. (#149)

Contributors

Thank you to @RobertSigmundsson for #147, #148, and #149, all found while migrating a production install onto 3.0.3.

Full diff: v3.0.3...v3.1.0