Skip to content

v1.8.0 - Indexer Durability & Resilience

Choose a tag to compare

@solaitken solaitken released this 14 Jun 05:00
· 107 commits to main since this release
7cdbfc0

Open Second Brain v1.8.0 - Indexer Durability & Resilience

Interrupting a long index run no longer risks losing work or wedging the index. A signal to o2b search watch now drains the in-flight pass at a file boundary and flushes before exiting instead of killing it mid-write; cancellation is cooperative through an AbortSignal composed into the existing Safeguard; a full reindexVault rebuild becomes resumable behind an opt-in flag, picking up a compatible staging build instead of starting over; and a writer-lock heartbeat plus a WAL-flush-on-exit registry keep a long run from looking stale and keep a bypassed close from leaving an orphan WAL. The suite reuses the existing safeguard and per-vault lock rather than adding a lifecycle subsystem, and every behaviour defaults to today's - a vault that sets nothing is byte-identical in results, ordering, and shape.

How the Indexer Durability suite keeps interrupted runs safe

What ships

  • No mid-write kill (o2b search watch). SIGINT/SIGTERM stops accepting new flushes, aborts the in-flight pass at its next file boundary, and awaits it to settle before exiting - bounded by search_shutdown_grace_seconds (default 5; 0 exits immediately after signalling). indexInto closes its store in a finally, so the aborted pass still consolidates the WAL and releases the writer lock. A second signal falls back to the default terminate. The flush/shutdown coordination lives in a testable IndexWatchRunner.
  • Cooperative abort (Safeguard + AbortSignal). The cooperative deadline gains an optional AbortSignal: one checkpoint() trips on either an aborted signal (a new SafeguardAbortError, checked first) or the existing timeout. The signal threads through indexVault and populateEmbeddings, checked at the same boundaries the deadline uses - between files and between embed batches, never mid-write. Bun's SQLite is synchronous, so abort is cooperative, never preemptive; the deletion sweep runs only on full completion, so an aborted run leaves a consistent, partially-refreshed index.
  • Opt-in resumable reindex (search_resume_reindex). An interrupted full rebuild no longer discards all progress: a compatible in-progress brain.sqlite.new staging build is resumed via the incremental fastpath instead of rebuilt from scratch. Resume is gated on a signature marker (schema version + chunk parameters + embedding signature) stored in the staging DB's index_state KV - no schema migration - so a drifted or unreadable staging DB is discarded and rebuilt, never trusted. The marker is cleared before the atomic swap, so the live index never carries staging state. Default off keeps the always-fresh rebuild.
  • Writer-lock heartbeat + WAL-flush-on-exit. The async writer lock refreshes its mtime mid-run (an explicit heartbeat below the 60s stale window) so a long index is never mistaken for a stale lock. A process-exit registry consolidates each open writer's WAL on a bypassed close(), mirroring the existing sync-lock cleanup hook.

Process wins

  • Every behavioural change defaults to today's behaviour. With search_resume_reindex off and no abort/grace configured, the index path is byte-identical to before - no staging marker is written, no extra open/close cycles. The two new config keys both default to current behaviour.
  • The incremental index path was already resumable (the mtime+size fastpath skips committed files; populateEmbeddings only computes missing vectors), so no redundant checkpoint mechanism was added for it. The genuinely non-resumable path - a full reindex - is the only one that gained a resume, and only behind a flag.
  • Honest multi-instance story, no fabricated daemon: the MCP server is stdio-only, so isolation comes from the per-dbPath writer lock. Two instances on different vaults run conflict-free; a second writer on the same vault gets a typed INDEX_LOCKED. Pinned by tests; no --port/--instance model was invented for an architecture that has no port.
  • Quality record: 4,542 tests / 0 fail, TypeScript clean, lint at 0 errors, version synced across every manifest, one CodeRabbit pass with no actionable findings.

Notes

  • The version bump to 1.8.0 shipped inside the feature PR (#97), per the project rule in CLAUDE.md.
  • Cancellation is cooperative by construction: Bun runs SQLite synchronously, so a run is stopped at its next natural boundary, never preempted mid-write. The release does not claim otherwise.
  • Release image: the canonical terminal style (animated GIF in this body; static PNG and the SVG source attached as assets).