v3.5.1 — smem consolidate stops exhausting the socket table
A maintenance release built around one symptom: smem consolidate stalling for minutes, flooding
the terminal with [Errno 104] Connection reset by peer, and ending in a report of zeros that
looked exactly like "there was nothing to do". Measuring it turned up several independent defects
feeding that one symptom, plus three unrelated fixes that landed in the same window.
Upgrade notes
- An
http(s)://SURREALDB_URLis now rewritten tows(s)://. HTTP is no longer a selectable
transport for the SurrealDB connection. This is deliberate: the SDK's HTTP transport opens a new
TCP connection per RPC (measured: two sockets per query, versus zero over WebSocket), which is
what made the reported failure reachable at all. An explicitws://,wss://or embedded
(memory:/surrealkv:) URL passes through untouched. smem consolidatenow exits non-zero when a stage failed. Previously a raising strategy
crashed the command; now the pass survives, the remaining strategies still run, and the summary
names the casualty — so the exit code has to carry the failure for anything that gates on it.- No schema change, no migration, no configuration change required.
Fixed — the consolidation connection storm
- Port exhaustion on the HTTP transport. Consolidation issues tens of thousands of small
queries (compressruns one per neuron id;semantic_linkwrites per row). With a fresh TCP
connection per RPC that exhausts the ephemeral port range within minutes — measured 39,720
sockets inTIME_WAITagainst the DB port — after which the kernel resets every new connection,
including the SDK's own reconnect and signin, and the client spins in[Errno 104]forever.
The store now multiplexes every RPC over one persistent WebSocket connection. - One dropped connection no longer costs one reconnect per concurrent caller. The re-auth lock
serialised concurrent callers but did not deduplicate them, so every query in a batch fan-out
that failed on the same dead connection built its own replacement, retried three times, and
logged its own warning. Reconnects are single-flight behind a generation counter now: the first
caller repairs the connection, the rest re-run their query on it, and only the path that actually
reconnected warns. The replaced connection is closed — it used to leak a socket per reconnect,
and on the WebSocket transport an orphaned receive task with it — and the backoff carries jitter. - A reconnect no longer cancels its siblings' in-flight queries. Closing a WebSocket connection
cancels every RPC future still pending on it, not just the one that triggered the close, and
asyncio.CancelledErroris aBaseException— so it walked past the retry path and tore down
whole batches. A cancellation the task did not request is treated as a dropped transport and
retried; a genuine cancellation still propagates untouched. - A storage instance reused from a second event loop repairs itself. The SDK creates its
response futures on the loop that opened the socket, so a cached storage singleton driven by a
laterasyncio.run()failed every query withRuntimeError: … Future … attached to a different loop— neither an auth nor a transport error, so nothing retried it and the connection stayed
broken for the life of the process. initialize()survives a transient reset across the whole connect-and-prepare window. Signin,
use, the version gate, schema apply and migrations now re-run as one idempotent unit on a fresh
connection after a dropped transport. Credential errors and version-gate rejections still fail
fast — neither fixes itself by reconnecting.
Fixed — consolidation correctness and honesty
- Consolidation stops fetching embedding vectors it never reads.
find_neuronsincludes them by
default, sodreampulled 10,000 of them in a single response and bothinterferencescans
carried one per row across the whole brain. semantic_linkno longer scans the whole neuron table, nor holds the event loop. It pages per
neuron type so the filter runs in the database index, and the similarity pass yields periodically.
Holding the loop for minutes was long enough to miss the WebSocket keepalive, after which the peer
drops the connection.- The numpy-less similarity fallback is bounded, yields, and says so. Without numpy the pass fell
into a pure-python O(n²·d) double loop that cannot finish at the shipped caps and that the
per-strategy timeout cannot interrupt, because the loop never awaits — so it hung rather than
completing. It now processes a bounded slice and logs a warning naming numpy when it truncates. - The consolidation report says what failed. A strategy raising anything other than a timeout
used to abort the whole pass. Failures are recorded per strategy, the remaining strategies still
run, and both failures and timeouts are named in the summary — the latter were being recorded but
never printed. A clean run prints neither. mergemust not absorb a pinned fiber. The merged fiber keeps onlymerged_fromand drops
source metadata, so a fiber pinned for any reason other than a habit/reasoning marker was absorbed
on the next pass and the pin disappeared with it. Pinning exists to make a fiber survive
lifecycle, and merge is part of lifecycle.
Fixed — data integrity
smem_editrefreshescontent_hashand the embedding on a content change. A content edit
replaced the text and nothing else, and because the storage read surfaces the stored vector back
intometadata["_embedding"], the old vector was actively re-saved against the new text. The
memory stayed retrievable by what it used to say — worst case, an entry edited specifically to
say it is outdated, still returned by semantic recall as current.reindex --missing-onlycould
not see the damage because the vector field was never empty; only a full--allre-embed repaired
it.content_hashdrifted the same way, so near-duplicate detection compared fingerprints of text
that no longer existed. Both derived fields are refreshed now; if the embedding provider is
unavailable the edit still succeeds and the stale vector is reported with a warning.
Fixed — query results and the hub API
- The
SELECT VALUEunwrap heuristic is gone.result[0] if isinstance(result[0], list) else resultwas a fossil from SDK 1.x, which returned the RPC envelope; the pinned
surrealdb>=2.0.0,<3.0.0unwraps it itself. Its only remaining effect was corruption: a
SELECT VALUEon an array-typed field returns a list of per-row arrays, and the heuristic
collapsed that to the first row's array. The retry/reconnect core moved to_query_response, with
_query(rows) and_query_values(per-row values) as honest thin wrappers, and new live
integration tests pin all three result shapes. GET /hub/status/{brain_id}andGET /hub/devices/{brain_id}bound the id length. The path
parameters were validated by character-set regex alone — an unbounded+— so a charset-valid but
over-long id cleared validation and failed later in the storage layer as a 500. The same id was
already a clean 422 on POST, where the Pydantic bound applied; only the path-param route bypassed
it. Both validators check length before the pattern now.
Verification
- Reproducing CI's
Integration (SurrealDB)job locally against a throwaway container, using that
job's own URL shape and command: the release base ran7110 passed, 5 errorswith 12Errno 104
lines in 133.8 s; this release runs7172 passed, 0 errorswith 0Errno 104lines in 71.7 s.
The halved wall clock is the same effect from the other side — half the time was going into
opening and tearing down two TCP connections per RPC. - An end-to-end consolidation pass against a real database went from hundreds of failed reconnect
attempts and a socket table driven past the ephemeral port range to zero reconnect attempts
and no measurableTIME_WAITgrowth, with the peak event-loop stall dropping well under the
WebSocket keepalive budget. - Full quality gate green (lint, format, mypy, tests at the 67 % coverage floor, security scan);
the suite was additionally run over WebSocket, over an HTTP-shaped URL, with thestressmarker,
and repeated for order stability.
Commits in this release
4b133af1fix: stopsmem consolidateexhausting the socket table on the HTTP transport (#171)bdb57a95fix:initialize()retries the whole connect-and-prepare window, not just signin (#173)fef4fb97fix: remove theSELECT VALUEunwrap heuristic — SDK 2.x already unwraps (#169)14958889fix(mcp):smem_editrefreshescontent_hashand the embedding on a content change (#166)be4c9e21fix(consolidation): merge must not absorb a pinned fiber (#168)96837744fix:initialize()retries a transient connection reset during signin (#172)64548880fix(hub): boundbrain_id/device_idlength so over-long ids 422 instead of 500 (#167)
Full Changelog: v3.5.0...v3.5.1