Repository navigation
Releases: Pradyothsp/GoVec
Release list
v0.2.1
Better default recall on text embeddings, and a fix so ef_search changes reach existing data.
Changes
ef_searchnow defaults to 100 (was 50). Recall@10 on 1536-d text embeddings goes from 0.97 to 0.995, at a bit more query latency.ef_constructionnow defaults to 100 (was 200). Inserts are faster with the same recall at 10k vectors. On larger datasets, set it back to 200 if you need the extra recall.- Fixed: a snapshot overrode the configured
ef_searchon load, so changing it had no effect on data saved earlier.
No format or API changes; 0.2.0 data loads as is.
The README has new benchmark numbers on 100k text embeddings, and the roadmap lists what's next.
docker run -p 9697:9697 -v govec-data:/data ghcr.io/pradyothsp/govec:0.2.1Full Changelog: v0.2.0...v0.2.1
v0.2.0
Durability fixes, a 3x faster search at OpenAI embedding size, a binary WAL, and a stricter, more consistent search API.
Before you upgrade
- Data files from 0.1.0 can't be read. The snapshot format and the WAL format both changed (below). A 0.2.0 server pointed at 0.1.0 data refuses to start and says why. There is no migration: start with an empty data directory and insert your vectors again.
kis now required on search, over REST and gRPC. A REST request withoutkused to be accepted; it now returns 400.k=0returns an empty list on both index types. Brute force used to return every vector.- A negative
kis rejected with 400 /InvalidArgument. It used to crash HNSW search. - An empty
sparse_vectormeans none on both transports. REST used to answer it with 400.
Use the Python SDK 0.1.2, which mirrors the new k field. 0.1.1 also works, since it always sends k, except for one edge: over gRPC its k=0 arrives as "missing" and is rejected rather than answered with an empty list.
Durability
- A failed snapshot no longer loses writes. Every save truncated the WAL before the new snapshot was safely on disk. Now the snapshot is fsynced, renamed and its directory fsynced before the WAL is truncated.
- Large snapshots no longer run out of memory. Snapshots are streamed instead of built whole in memory: 100k vectors at 1536 dimensions used to be OOM-killed mid-save. HNSW snapshots store the graph as topology only, so each vector is written once; a snapshot is now about 1.03x the raw vectors, down from about 2x.
- Reset survives a restart. It used to clear memory and the WAL but keep the snapshot, so a restart brought the vectors back, and vectors of a new width inserted after the reset stopped the server from starting. Reset now publishes an empty snapshot first; a failed reset changes nothing and returns 500 /
Internal. - A rejected insert can't block the next restart. A vector refused for its width, a bad sparse vector or an out-of-range int8 value was written to the WAL anyway, and replay stops at a record it can't apply. Writes are now checked before they are logged.
- A restart remembers the learned dimension. Without it, the first insert after a restart set the width, and one wrong-width vector broke the index.
Performance
- About 3x faster search at 1536 dimensions. The float32 cosine distance added every element into one running total, so each add waited on the last; it was 84% of search time. Eight independent totals, same float64 precision: 1.73 → 0.38 µs per comparison, and in-process HNSW search on 10k OpenAI embeddings went from 3.38 to 1.04 ms.
- Faster brute-force and Euclidean distance: 3.7–5.4x at 1536 dimensions. The squared Euclidean kernel is now shared by both index types.
- Binary WAL: about 3x smaller (1.01x the raw vector, from 19 KB per 1536-d entry in JSON) and about 45x faster to replay. Every record carries a CRC-32C; a damaged record is treated as a torn write only if nothing valid follows it.
- gRPC streaming batch insert now goes to the engine in chunks of 1000, with one WAL fsync per chunk instead of one per vector. If a chunk fails, its items are reported as per-item errors and
inserted_countstays exact. - Exact-size vector storage: each float32 vector is stored in its own exact-size slice instead of the decoder's, which carried about 17% spare capacity at 1536 dimensions.
Fixes
- HNSW's
/infonow reportsenable_mmap(it was always false). - HNSW no longer learns the index width from a vector it then rejects.
- A rejected insert no longer leaves metadata or sparse-vector index entries behind (both index types).
- Sparse vectors are validated once, in the engine: the same rule and the same 400 /
InvalidArgumentmessage on both transports. config.example.yamlnow matches the real defaults.
Run it
docker run -p 9697:9697 -v govec-data:/data ghcr.io/pradyothsp/govec:0.2.0Internal
Both index types now share one implementation of storage, WAL ordering, replay, reset, scoring and snapshots, so a rule is written once. The test suite was consolidated (duplicates removed, engine pairs and one-case tests turned into tables, Arrange / Act / Assert throughout) with per-package coverage held or raised. The Docker build context is an allowlist of the server's sources. A security policy points at private vulnerability reporting.
Full Changelog: v0.1.0...v0.2.0
v0.1.0
First public release of GoVec: a compact vector search engine in Go, served over REST and gRPC from a single static binary.
Highlights
- HNSW with heuristic neighbour selection, plus exact brute-force search; cosine and Euclidean distance
- Scalar int8 quantization: -61% disk and -26% RAM at 100k vectors
- Hybrid dense + sparse search and metadata filtering
- Write-ahead log and atomic snapshots, with crash recovery on startup
- REST and gRPC with identical semantics, including streaming batch insert over gRPC
- Multi-arch container image (linux/amd64, linux/arm64)
Run it
docker run -p 9697:9697 -v govec-data:/data ghcr.io/pradyothsp/govec:0.1.0REST listens on 9697; gRPC on 9698 when grpc.enabled is true. See the README for the API and configuration.
Known limitations
- Single node, in-memory first
- Filtered HNSW search can return fewer than
kresults for selective filters - Hybrid search scores sparse vectors with an O(n) scan
Related
- Python SDK: govec-python
- Benchmarks against Chroma and Qdrant: govec-bench
Full Changelog: https://github.com/Pradyothsp/govec/commits/v0.1.0