Skip to content

Benchmark ‐ stable‐PoC

klmckeig edited this page Aug 10, 2026 · 7 revisions

Branch: stable-PoC

Numbers on this page reflect the PoC build. Final released performance may differ somewhat due to productization overhead. See later benchmark pages in this series for updated results as the extension matures.

Summary

The SVS extension adds Vamana, an open-source graph-based ANN index, to PostgreSQL as a new pluggable index type layered on top of pgvector (same data types, same distance metrics, no fork of pgvector itself). Optionally, the LeanVec and LVQ compression layer can be enabled for further memory and performance gains.

This page benchmarks three configurations against each other:

  • HNSW — pgvector's existing baseline index
  • SVS Vamana — the open-source core, no compression
  • SVS Vamana + LeanVec — Vamana with the optional compression layer enabled

All configurations were tuned to comparable recall (~95%) so that QPS and latency numbers are measured on equal footing.

Benchmark Configuration

Benchmark tool VectorDBBench (forked to add support for the Vamana index type)
Server r8i.4xlarge
Client r8i.8xlarge
OS Amazon Linux 2023
Precision fp32
Datasets Cohere 1M (768 dimensions), OpenAI 500K (1536 dimensions)
search_num_threads 13

Index/search parameters:

Cohere 1M (768D) HNSW SVS Vamana SVS Vamana + LeanVec
k 100 100 100
ef_search / search_window_size 170 237 162
M / graph_degree 16 32 32
build_window_size 200 200 200
recall 95.03 95.04 95.04
OpenAI 500K (1536D) HNSW SVS Vamana SVS Vamana + LeanVec
k 100 100 100
ef_search / search_window_size 118 119 200
M / graph_degree 16 32 32
build_window_size 200 200 200
recall 95.04 95.02 95.04

Results: Cohere 1M (768D)

Concurrency HNSW QPS HNSW p99 (ms) SVS Vamana QPS SVS Vamana p99 (ms) SVS Vamana + LeanVec QPS SVS Vamana + LeanVec p99 (ms) Vamana QPS vs. HNSW Vamana+LeanVec QPS vs. HNSW Vamana latency improvement Vamana+LeanVec latency improvement
5 1,100 6.2 1,493 4.0 4,175 1.9 1.36x 3.80x 35% 69%
10 2,197 6.3 2,728 5.3 7,821 1.9 1.24x 3.56x 16% 70%
20 3,319 9.7 3,564 7.4 11,515 3.3 1.07x 3.47x 24% 66%
30 3,375 16.7 3,506 12.6 10,824 8.7 1.04x 3.21x 25% 48%
40 3,356 26.0 3,870 15.8 12,347 11.2 1.15x 3.68x 39% 57%
50 3,328 35.9 3,988 18.8 12,178 14.6 1.20x 3.66x 48% 59%
image image

Peak result: up to 3.80x QPS (at concurrency 5) and up to 70% latency improvement (at concurrency 10) for SVS Vamana + LeanVec vs. HNSW.

Bare Vamana, with no compression at all, also outperforms HNSW at every concurrency level tested on this dataset (1.04x–1.36x QPS), demonstrating that the open-source core alone is competitive before any compression is applied.

Results: OpenAI 500K (1536D)

Concurrency HNSW QPS HNSW p99 (ms) SVS Vamana QPS SVS Vamana p99 (ms) SVS Vamana + LeanVec QPS SVS Vamana + LeanVec p99 (ms) Vamana QPS vs. HNSW Vamana+LeanVec QPS vs. HNSW Vamana latency improvement Vamana+LeanVec latency improvement
5 1,174 6.0 1,128 4.0 3,166 2.4 0.96x 2.70x 33% 60%
10 2,298 6.6 2,087 5.3 6,060 2.4 0.91x 2.64x 20% 64%
20 3,530 9.7 2,601 7.4 8,518 3.5 0.74x 2.41x 24% 64%
30 3,250 17.6 2,438 12.6 7,610 5.3 0.75x 2.34x 28% 70%
40 3,179 28.8 2,635 15.8 8,096 6.8 0.83x 2.55x 45% 76%
50 3,061 43.5 2,684 18.8 8,753 8.0 0.88x 2.86x 57% 82%
image image

Peak result: up to 2.86x QPS and up to 82% latency improvement (both at concurrency 50) for SVS Vamana + LeanVec vs. HNSW.

On this higher-dimensional dataset, bare Vamana without compression trails HNSW slightly on QPS (0.74x–0.96x) at higher concurrency, while still improving p99 latency. This is a useful data point on its own: at 1536 dimensions, the LeanVec compression layer is what delivers the throughput win, not the Vamana graph alone. Both configurations are shown here so readers can evaluate the open-source core and the optional compression layer independently.

Observations

  • The open-source core holds up on its own. Without any compression, Vamana is competitive with HNSW on both datasets: ahead on Cohere 1M (768D), slightly behind on OpenAI 500K (1536D) at higher concurrency.
  • The optional compression layer (LeanVec) delivers the larger gains, especially at higher dimensionality, where it's the difference between trailing HNSW and leading it by up to 2.86x.
  • Recall was held constant (~95%) across all configurations, so these QPS/latency comparisons are on equal footing, not an artifact of looser recall targets.

Questions or feedback on this benchmark? Open an issue or discussion in this repository.

Clone this wiki locally