2.11.0 — Fast bulk doc-training
smem train on large documents is dramatically faster on big brains, and the encoder now batches its writes instead of one round-trip per neuron/synapse.
Added
add_synapses_batch(multi-statement INSERT RELATION) + base default with per-synapse fallbackfind_neurons_exact_batchoverride — onecontent IN $contentsround-trip for N lookups (was N+1)- tqdm progress bar in
smem train(optional import +logger.infofallback every 50 chunks)
Changed
increment_keyword_df: per-keyword SELECT+merge N+1 → one multi-statement UPSERT (#1 per-chunk op count, ~93/chunk → 1)CreateSynapsesStep+CoOccurrenceSteppersist viaadd_synapses_batchthrough_persist_synapseshelperfind_neurons:brain_idinlined as a literal instead of$brain_id— SurrealDB 3.2.0 only uses the brain_id index for an inline literal; a parameterized value full-scans. EXPLAIN:TableScan→IndexScan [idx_neuron_brain]
Performance
- In-memory SDB 3.2.0, N=100: 1.446 → 0.847 s/chunk (-41%, <1s target)
- ~10× fewer DB ops/chunk (581 → 58) on a full 6651-chunk run
find_neuronsindex-driven on big brains — the dominant lever on disk-backed brains
Full changelog: https://github.com/acidkill/surreal-memory/blob/main/CHANGELOG.md#2110--fast-bulk-doc-training