Skip to content

2.11.0 — Fast bulk doc-training

Choose a tag to compare

@acidkill acidkill released this 17 Jul 10:16
· 170 commits to main since this release
ec48d93

smem train on large documents is dramatically faster on big brains, and the encoder now batches its writes instead of one round-trip per neuron/synapse.

Added

  • add_synapses_batch (multi-statement INSERT RELATION) + base default with per-synapse fallback
  • find_neurons_exact_batch override — one content IN $contents round-trip for N lookups (was N+1)
  • tqdm progress bar in smem train (optional import + logger.info fallback every 50 chunks)

Changed

  • increment_keyword_df: per-keyword SELECT+merge N+1 → one multi-statement UPSERT (#1 per-chunk op count, ~93/chunk → 1)
  • CreateSynapsesStep + CoOccurrenceStep persist via add_synapses_batch through _persist_synapses helper
  • find_neurons: brain_id inlined as a literal instead of $brain_id — SurrealDB 3.2.0 only uses the brain_id index for an inline literal; a parameterized value full-scans. EXPLAIN: TableScan → IndexScan [idx_neuron_brain]

Performance

  • In-memory SDB 3.2.0, N=100: 1.446 → 0.847 s/chunk (-41%, <1s target)
  • ~10× fewer DB ops/chunk (581 → 58) on a full 6651-chunk run
  • find_neurons index-driven on big brains — the dominant lever on disk-backed brains

Full changelog: https://github.com/acidkill/surreal-memory/blob/main/CHANGELOG.md#2110--fast-bulk-doc-training