Repository navigation
v1.8.0 - Indexer Durability & Resilience
Open Second Brain v1.8.0 - Indexer Durability & Resilience
Interrupting a long index run no longer risks losing work or wedging the index. A signal to o2b search watch now drains the in-flight pass at a file boundary and flushes before exiting instead of killing it mid-write; cancellation is cooperative through an AbortSignal composed into the existing Safeguard; a full reindexVault rebuild becomes resumable behind an opt-in flag, picking up a compatible staging build instead of starting over; and a writer-lock heartbeat plus a WAL-flush-on-exit registry keep a long run from looking stale and keep a bypassed close from leaving an orphan WAL. The suite reuses the existing safeguard and per-vault lock rather than adding a lifecycle subsystem, and every behaviour defaults to today's - a vault that sets nothing is byte-identical in results, ordering, and shape.
What ships
- No mid-write kill (
o2b search watch). SIGINT/SIGTERM stops accepting new flushes, aborts the in-flight pass at its next file boundary, and awaits it to settle before exiting - bounded bysearch_shutdown_grace_seconds(default 5;0exits immediately after signalling).indexIntocloses its store in afinally, so the aborted pass still consolidates the WAL and releases the writer lock. A second signal falls back to the default terminate. The flush/shutdown coordination lives in a testableIndexWatchRunner. - Cooperative abort (
Safeguard+AbortSignal). The cooperative deadline gains an optionalAbortSignal: onecheckpoint()trips on either an aborted signal (a newSafeguardAbortError, checked first) or the existing timeout. The signal threads throughindexVaultandpopulateEmbeddings, checked at the same boundaries the deadline uses - between files and between embed batches, never mid-write. Bun's SQLite is synchronous, so abort is cooperative, never preemptive; the deletion sweep runs only on full completion, so an aborted run leaves a consistent, partially-refreshed index. - Opt-in resumable reindex (
search_resume_reindex). An interrupted full rebuild no longer discards all progress: a compatible in-progressbrain.sqlite.newstaging build is resumed via the incremental fastpath instead of rebuilt from scratch. Resume is gated on a signature marker (schema version + chunk parameters + embedding signature) stored in the staging DB'sindex_stateKV - no schema migration - so a drifted or unreadable staging DB is discarded and rebuilt, never trusted. The marker is cleared before the atomic swap, so the live index never carries staging state. Default off keeps the always-fresh rebuild. - Writer-lock heartbeat + WAL-flush-on-exit. The async writer lock refreshes its mtime mid-run (an explicit heartbeat below the 60s stale window) so a long index is never mistaken for a stale lock. A process-exit registry consolidates each open writer's WAL on a bypassed
close(), mirroring the existing sync-lock cleanup hook.
Process wins
- Every behavioural change defaults to today's behaviour. With
search_resume_reindexoff and no abort/grace configured, the index path is byte-identical to before - no staging marker is written, no extra open/close cycles. The two new config keys both default to current behaviour. - The incremental index path was already resumable (the mtime+size fastpath skips committed files;
populateEmbeddingsonly computes missing vectors), so no redundant checkpoint mechanism was added for it. The genuinely non-resumable path - a full reindex - is the only one that gained a resume, and only behind a flag. - Honest multi-instance story, no fabricated daemon: the MCP server is stdio-only, so isolation comes from the per-
dbPathwriter lock. Two instances on different vaults run conflict-free; a second writer on the same vault gets a typedINDEX_LOCKED. Pinned by tests; no--port/--instancemodel was invented for an architecture that has no port. - Quality record: 4,542 tests / 0 fail, TypeScript clean, lint at 0 errors, version synced across every manifest, one CodeRabbit pass with no actionable findings.
Notes
- The version bump to 1.8.0 shipped inside the feature PR (#97), per the project rule in
CLAUDE.md. - Cancellation is cooperative by construction: Bun runs SQLite synchronously, so a run is stopped at its next natural boundary, never preempted mid-write. The release does not claim otherwise.
- Release image: the canonical terminal style (animated GIF in this body; static PNG and the SVG source attached as assets).
