Releases: runcaptain/compass
Release list
v0.4.0 — Serverless: stateless writers, tenant partitions, serve-from-storage
Compass v0.4.0 — Serverless: stateless writers, tenant partitions, serve-from-storage
v0.3.0 made object storage the source of truth. v0.4.0 makes Compass a
serverless database: nodes are disposable, tenants are cheap, and a cold
collection answers its first query in milliseconds — while local-first,
zero-dependency embedded mode remains byte-for-byte the default.
Warm serverless (roadmap Phases 0–3)
- Stateless writer role (
COMPASS_ROLE=writer): append-only nodes with no
local indexes and instant boot. Durable immediately; searchable on serving
nodes within the refresh interval. - CAS id-block allocator: attached nodes and writers can never mint
colliding chunk ids. - Bucket collection config: vector-space specs survive cold rebuilds.
- Manifest refresher + read-your-writes: serving nodes converge on other
nodes' writes (default 5s); writes returnseq;min_seqon search gives
read-your-writes with a bounded wait. - Lazy attach + LRU detach (
COMPASS_LAZY_ATTACH,COMPASS_MAX_ATTACHED).
Tenant partitions (Phase 6)
config.partition_by = "tenant_id"at create: every chunk routes to an
internal per-tenant partition — a full engine namespace behind one
collection API. Searches/deletes filter by the partition field (set
membership fans out ≤16, merged); ids stay collection-unique; partitions
auto-create on first ingest (writer role included), attach on demand, hide
from listings, cascade-delete with the parent.- Measured: 50–200 tenants in one collection, boot 0.6s / ~15MiB RSS
regardless of tenant count; serving RAM tracks the hot-tenant set. - Works fully in LOCAL mode too (partition dirs on disk, shared id file).
Serve-from-storage (Phase 5) — true serverless
COMPASS_COLD_SERVE=true: semantic queries on UNATTACHED collections and
partitions are answered straight from object storage — manifest read,
cached centroids, a few cluster range-GETs, byte-range hydration.- Segment format CSEG0003: IVF-clustered vectors + row-addressable metadata,
built at compaction. v2 segments stay readable; pre-v0.4 readers fail loudly
on v3 (do not mix writer versions against one bucket). - Measured (300k × 64d, MinIO): fresh node boots 0.3s at 28MiB; FIRST query
70ms; steady state p50 25ms / p95 37ms at 66MiB RSS — vs minutes and ~2GiB
for the attach path. - Cold reads see the full committed state including the WAL tail —
read-your-writes by construction. Metadata filters apply. FTS on a cold
namespace returns a clear error (attach still builds full local indexes). COMPASS_WARM_AFTER(default 3) promotes hot namespaces to a background
attach: cold → warm → hot automatically.
Fixed
- Facets: wiped by every ingest after the first, empty after restart, counted
deleted chunks (latent since v0.2) — now chunk-id-keyed roaring treemaps. - Warm restarts were O(collection) (full HNSW rebuild on a legitimately-stale
index file): now incremental heal; 100k restart 20.2s → 1.6s. - Vector-space rebuild completion never activated the space or hot-loaded the
index until restart. - Missing collections returned HTTP 500/400; typed not-found now maps to 404.
- Sub-1000-vector keymap persistence (wrong ids under non-dense allocation).
Lean pass
- ~1,200 lines of dead weight removed (unwired GPU backend plumbing, dead
filter evaluators, never-wired persistence codecs, rayon); dead-code lint
re-enabled crate-wide; one filter semantics (roaring pushdown) for search
AND delete-by-filter; collections/mod.rs halved.
Compatibility & migration
- Local mode: no changes required; no cloud features activate without cloud
config (guard-tested). - Cloud mode: pre-v0.4 namespaces migrate organically (bucket config
back-filled, id allocator seeded from the high-water mark). Do NOT run
v0.3 and v0.4 writers against the same bucket during a rolling upgrade;
v0.3 readers cannot read v0.4 (CSEG0003) segments and fail loudly. - New envs: COMPASS_ROLE, COMPASS_REFRESH_INTERVAL, COMPASS_LAZY_ATTACH,
COMPASS_MAX_ATTACHED, COMPASS_COLD_SERVE, COMPASS_WARM_AFTER,
COMPASS_COLD_NPROBE, COMPASS_MAX_CONCURRENCY (all documented in
.env.example; all optional).
Evidence
- 106 local / 159 object-storage tests; clippy clean under -D warnings (lint
fully enabled); four build combinations verified in CI (including the
non-test cloud build that only Docker used to compile). - 58-check live E2E across full node + writer node + cold node on MinIO.
- Live benches: v0.3.0 comparison (+48% ingest, −57% RSS, −72% bucket bytes,
equal search), 200-tenant bounded-RAM proof, 300k cold-serve proof.
v0.3.0 — Object-Storage Persistence, Typed Relations, Deletes
Compass 0.3.0 turns the embedded engine into a cloud-durable one: your own object storage can now be the source of truth, while local-first stays the zero-config default.
Highlights
🪣 Object-storage persistence (opt-in) — build with --features object-storage and set COMPASS_STORAGE=s3://bucket[/prefix] (or gs://, az://). Every write lands durably in your bucket first — an LSM of immutable WAL fragments committed through a CAS'd manifest, with automatic compaction and physical GC. Destroy the container and its disk; on the next boot Compass discovers every collection in the bucket and rebuilds all local indexes — chunks, hierarchy, and relations included. Works with AWS S3, GCS, Azure Blob, MinIO, and R2; an optional key prefix lets multiple deployments share one bucket.
🔗 Typed chunk relations — directed, labeled, many-to-many edges between chunks (cites, supersedes, anything), disk-backed in redb with multimap endpoint indexes. Create in batches, query per chunk with direction/type filters, or return them inline with search hits. Dangling targets are reported (target_status: "missing"), never silent.
🗑️ Deletes + compaction — soft-delete by id or metadata filter (DELETE /chunks/:id, POST /delete); tombstoned chunks vanish from results immediately and are physically reclaimed by compaction (POST /compact, or automatic in cloud mode).
🔢 True u64 chunk ids — the filter index moved from RoaringBitmap to RoaringTreemap, and the mmap vector file gained a u64-count v2 header (legacy files read transparently and migrate on first append). No more silent truncation past 4.29B ids.
Hardening
This release went through three adversarial review rounds. Among the bugs found and regression-tested before release: concurrent-append data loss, local↔cloud split-brain in both directions, a compaction path that leaked objects unboundedly, chunk-id reuse after compaction, GCS conditional writes requiring the object generation rather than the ETag, silently-dropped bucket prefixes, torn mmap appends, and missing embedding-dims validation. Details in the CHANGELOG.
Testing
99 tests in the default build, 125 with object-storage — including fault injection, 16-writer concurrent appends, torn-append recovery, and restart recovery on both fresh and persistent disks. Env-gated integration tests run the CAS and LSM lifecycle against a real S3 endpoint (verified on MinIO): see crates/compass/src/storage/s3_integration_tests.rs for the recipe, or docker compose -f docker-compose.minio.yml up to try it locally.
Scope, honestly
Cloud mode is durable via object storage, fast via local: reads serve from locally rebuilt indexes, so this is not stateless multi-node serving yet, and cold-start recovery materializes a collection's live set in RAM. Those are the next frontier — see the CHANGELOG's scope notes.
Full changelog: v0.2.0...v0.3.0
v0.2.0 — Scalability + Open-Source Readiness
Scalability
- Memory-mapped vector storage (
MmapVectors). ReplacesVec<Vec<f32>>with a flat mmap-backed file. Zero-copy reads, append-only writes, near-zero RSS for vector data regardless of dataset size. - Disk-backed chunk metadata (
ChunkStore). Backed by redb (pure Rust embedded DB). Replaces in-memoryHashMap<u64, DocumentChunk>. - Incremental HNSW indexing. Ingest path loads existing USearch index via
.load(), appends with.add(), saves — no more full rebuild on every ingest. - New dependencies:
memmap2,bytemuck,redb
Telemetry
- Anonymous opt-out telemetry via PostHog. Sends daily heartbeat with instance UUID, version, OS, arch, collection count, total vectors. No data content or PII.
- Opt out:
COMPASS_TELEMETRY=offorDO_NOT_TRACK=1 - Python examples (
examples/python/compass_client.py) covering all API interactions.
Open-Source Readiness
CODE_OF_CONDUCT.md(Contributor Covenant v2.1)SECURITY.md(vulnerability disclosure policy).github/ISSUE_TEMPLATE/(bug report, feature request).github/PULL_REQUEST_TEMPLATE.md.github/CODEOWNERS.github/workflows/release.yml(automated releases on tag push)docker-compose.yml(local quickstart)- README badges (CI, license, Rust version)
- Compass logo + Captain wordmark in README header
Changed
- MSRV bumped to 1.88 (from 1.82). Required by
time@0.3.47. - Dockerfile upgraded to
rust:latest+debian:trixie-slim.
v0.1.0 — Initial Release
Initial Release
Compass is an embedded vector + full-text search engine. Single binary, zero external dependencies.
Added
- Tantivy BM25 full-text search with bitset-faceted metadata
- USearch HNSW vector search (mmap-backed, disk-persistent)
- Reciprocal Rank Fusion hybrid search (k=60)
- Named vector spaces — multiple embedding models per collection
- One-click model upgrades with background re-embedding and atomic swap
- Parent-child relationships with TAMS-compatible hierarchies
- Query-time scoring: recency decay, metadata boost, relationship boost
- Native query embedding via Candle BGE-small (Rust, no Python)
- External GPU embedding endpoint support
Performance
- ~15k QPS per instance on 16-core hardware
- p99 < 50ms for top-10 retrieval
- ~1,500 docs/sec indexing with GPU-backed TEI
Quickstart
cargo build --release
./target/release/compass
# http://localhost:4001Or with Docker:
docker build -t compass .
docker run -p 4001:4001 -v ./data:/app/data compass