Skip to content

vector-store-1.10.0

Choose a tag to compare

@ewienik ewienik released this 07 Aug 16:25
· 188 commits to master since this release
e0d7b09

New Features

  • Local Vector Indexes on Non-Primary-Key Columns: Added support for local vector indexes whose partition key includes non-primary-key columns. Values for such columns are propagated through both the initial full scan and CDC, with the table layer synchronizing data from the two sources, and IndexMetadata was extended to carry primary-key columns and the partition key count. (#517)
  • Index Build Progress in Status Endpoint: The GET /api/v1/indexes/{keyspace}/{index}/status endpoint now reports a build_progress percentage derived from the initial full-scan progress, letting clients observe how far the backfill has advanced while an index is BOOTSTRAPPING. (#506)
  • Index Creation Options via HTTP API: Index creation options, previously visible only by querying system_schema.indexes via CQL, are now exposed through the HTTP API using data Vector Store already holds in memory. GET /indexes now includes each index's creation options, and a new GET /indexes/{keyspace}/{index} endpoint returns the same type and options information scoped to a single index. (#545)
  • Experimental In-Memory DiskANN Provider: Added an in-memory DiskANN index provider built on the new DiskANN DefaultProvider, with index parameters mapped from the index configuration. The DiskANN index type now supports adding and removing vectors, ANN search, count, partition removal, and memory checks before distributing messages. This integration is experimental. (#526, #540)

Performance

  • Search Prioritized Over Index Modifications: Fixed high search latency during massive vector deletions by splitting the index message stream into separate modify and search streams, with message selection biased toward search requests while avoiding starvation of the modification stream. (#542)
  • Interval-Based Tantivy Commits: FTS indexing now commits to Tantivy on a 3-second interval instead of per document, with a forced commit once 10,000 documents are pending, reducing indexing time for 10,000 small documents from roughly 2 minutes to under 1 second. This change also fixes an incorrect double reader reload after commits. (#521)

Configuration & Operations

  • CQL Preferred Datacenter/Rack: Added opt-in VECTOR_STORE_CQL_PREFERRED_DATACENTER and VECTOR_STORE_CQL_PREFERRED_RACK settings that route CQL reads to local nodes first, falling back to remote datacenters/racks when unavailable. This lets deployments that colocate Vector Store with a specific ScyllaDB datacenter or rack reduce cross-rack read traffic; setting the rack without the datacenter is rejected at config load time. (#546)

CI & Release

  • Release Workflow Fixes: The release workflow now uses the repository's VECTOR_GCP_* secrets for GCP authentication, fixing IAM access-token denials during releases, and the redundant build-cloud-images workflow that failed on every release was removed. (#539)
  • Docker Build Cache Permissions: Fixed Permission denied failures in the docker build job on fresh GitHub Actions runners by pre-creating the cargo cache directories before they are bind-mounted into the build container, so Docker no longer auto-creates them as root. (#537)

Internal Improvements

  • Single Worker Actor: Refactored the run function to pass memory and worker actors directly instead of through an index factory function, so a single worker actor instance can be shared by all indexes. Also added *.log files to .gitignore. (#536)
  • Rust 1.97.1: Updated the Rust toolchain to 1.97.1, which fixes a miscompilation in an LLVM optimization. (#544)

Documentation

  • Contributor Guidelines: Expanded CONTRIBUTING.md with documentation of the Validator end-to-end harness (building with the release toolchain and running against a ScyllaDB docker image), manual testing with the example docker-compose stacks, and a new Commit and PR Organization section based on the ScyllaDB patch organization guidelines, including Jira issue references. Also added a repository-level CLAUDE.md agent-instructions file surfacing these guidelines. (#533)
  • Docker Compose Examples Updated: Bumped the example docker-compose files to pin the latest ScyllaDB (2026.2.2) and Vector Store (1.9.1) images. (#529)

QA & Validator

  • Non-PK Local Index E2E Coverage: Added Validator end-to-end tests for the new local vector indexes based on non-primary-key columns, alongside unit tests for the table-layer changes. (#517)
  • Per-Row TTL on Already-Indexed Rows: Added a Validator test covering the "existing user" scenario where a per-row CQL TTL is added to rows that are already part of a fully built vector index, verifying the expiration propagates through CDC and removes the corresponding vectors from the index. (#532)
  • Large FTS Document Set Coverage: Added a Validator test that inserts 10,000 documents, waits for the FTS index to catch up, and asserts that a rare-term BM25 search returns only the expected documents and that a common-term search respects the query LIMIT, with indexing and query durations measured and logged. Existing FTS test cases were also refactored. (#521)
  • DiskANN Integration Tests: Refactored the vector-index integration tests into a common vs_index module parametrized by index configuration using rstest (fully replacing ntest), so each test now runs against both the usearch and DiskANN configurations. (#543)
  • Latte FTS Benchmark Workload: Added a Latte-based performance testing suite for ScyllaDB full-text search (BM25), covering the full lifecycle from data loading through index build to search, reporting indexing throughput and IR accuracy metrics (recall@k, precision@k, MRR, nDCG@k) against qrels, with a smoke-test fixture and README. (#520)
  • Search-and-Delete Benchmarks: Refactored the benches and added new cdc-delete and search-while-deleting performance tests for benchmarking parallel search and delete operations, plus a new delete-rows benchmark command to simulate massive deletions. Parquet dataset loading was also parallelized with uploading for faster benchmark runs. (#541)

New Contributors

Full Changelog: 1.9.1...1.10.0