Repository navigation
Releases: knowusuboaky/VectrixDB
Release list
VectrixDB 2.2.0
Added
- The server stays up under load.
- More MCP tools, resources and prompts.
- A skill for an assistant connected over MCP.
- The docs, grown.
- One release, every registry.
- The documentation's home page shows every way in:
- A client in four languages, one surface.
- Clients safe to point anywhere.
- A company's presets, its command's name, and its keychain.
- The
vectrixdbcommand on a server. - Run it inside a company's registry.
vectrixdb doctortries every part of an install.- A collection keeps up with feeds and pages.
- A release is one button and one approval.
- Container images, for Intel and ARM, signed.
vectrixdb extract-servestarts the extraction service- The key for an extraction service can come from a file.
vectrixdb checkknows the variables Kubernetes sets.- Searches, ingestion and evaluation runs can be traced, off until asked.
- MCP on the server, for a team or a company.
vectrixdb mcpover HTTP serves the reading tools only- A rechunk can be previewed.
- The docs have an index for coding agents and a prompt to give one.
- The test-question writer spreads its questions and judges harder.
stale_evidence()says which questions' quotes have left the index.- Moving pictures of the dashboard.
VectrixSync.cdc()carries deletes to the target.scripts/dashboard_shots.pyretakes the dashboard's pictures.- Single sign-on, before its provider is set up.
- A security group is narrowed to a list of people.
- Emergency sign-in.
- Single sign-on and the People list, together.
- Single sign-on needs somebody named.
- An app's token needs no place on the list.
- Developer Access.
- A path for each endpoint behind a gateway.
vectrixdb checksays what reads each kind of file.- The restricted note says why.
- The brand goes further.
- About, for admins.
- A NOTICE file travels with VectrixDB.
- A web page is read as its content.
- A PDF is read by PDFium, from where its text sits.
- A chart a PDF draws is described.
- The pages the rules are unsure of are read by sight.
- A file is known by its bytes.
- Spreadsheets read as they print.
- Word and PowerPoint read as they print.
- RTF and OpenDocument files are read.
- Who said what, in the languages it was said in.
- What a video shows.
- A picture's values come as rows.
- The walkthrough tries regions in the order you write them.
- The walkthrough makes a Translator.
- The steps have a requirements file.
Changed
vectrixdb checksays when a key leaves reads open.- The
mcpextra needs mcp 2 or later. - A YouTube video with captions is read from them, not transcribed.
- The authenticator's QR code is drawn as the dashboard's own.
- The walkthrough sets how the query app scales.
- Three ways to set sign-in, and passkeys with one of them.
- Emergency sign-in goes with every way in.
- Developer Access is found by the spinner.
- GraphRAG on Bedrock asks with Converse.
- The examples are not in the repository.
- Types are checked as Python 3.12
- The walkthrough's chat model is
gpt-5.4-mini. - The
documentsextra bringspypdfium2andxlrd, - Extraction quality judges prose.
- A masked JSON reply masks everything in it,
- A busy service answers 503.
Fixed
- A signed link stays secret.
- A bot check is refused, not indexed.
- Three optional models are found where they were published.
- Two results that score alike come back in the same order every time.
- Years and figures are not phone numbers.
- A table's stacked column headings head their own columns.
- A footnote mark glued to its word is cut loose on every page,
- Keyword and hybrid search run in the store that holds the text.
VECTRIXDB_OFFLINEholds for the speech model.- The audit before release.
- The README's pictures showed as broken images.
mypy vectrixdbfailed in CI with numpy 2.4's stubs.- The docs no longer describe what left with visibility.
- A guest opening a collection is not asked to sign in.
- Policied counts the same on both pages.
- A test of a killed writer miscounted.
- A fresh checkout passes its own tests.
- A collection's Overview drew nothing for somebody signed in.
- A tab the collection's policy refuses is no longer left blank.
- Counts on a deployment.
- The line under the single sign-on screens sits in the middle.
- The second vector installs with the Azure extra.
- The walkthrough's query app asks people to sign in.
- The walkthrough's query app finds its collections.
- The walkthrough's query app has guests.
- The walkthrough's Cosmos account has one database,
data_db. - The walkthrough's embedding model keeps up with an annual report.
- A collection kept in the cloud opens where nothing was written before.
- A Markdown table whose header leaves out the label column reads right.
- A web page's text has no empty lines of spaces and no empty items.
- An old page reads whole and in time.
- A web page is decoded in the charset it names.
- Step 03 no longer reads a working create as a refusal.
- The records' database is
data_db. - The walkthrough's app is deployed twice: reads and queries apart.
- Reading a file survives a bad hour.
- The walkthrough's dashboard shows the files that would not read.
- A run records the collections it ran over.
- A collection's knowledge graph is kept where every server reads it.
- The overview recomposed, and a guest's overview drawn from it.
- A collection's record is its policy, and nothing else.
Removed
- Visibility, the masking switch and per-document policies on records.
- The dashboard's charts, redrawn, and every list ten a page.
- The error catalogue.
- Masking with an engine that reads meaning, before a document is indexed.
- Who may retrieve from a collection: its policy, checked before anything is searched.
- Visibility is
privateorpublic, and a guest reads excerpts. - The dashboard shows and sets who may retrieve.
- The Azure walkthrough's app makes a collection the way its steps do.
- Each collection's rules in one record every server reads.
- The audit trail and the access log in append blobs nothing can rewrite.
- Readable Cosmos DB items.
- The chunking techniques compared, and the Evaluate page's Chunking tab.
- A cut and a way of searching a host can follow the runs for, or pin.
- Runs are kept by kind:
retrieval/runs/<id>/andchunking/runs/<id>/. - Four more ways to cut and embed.
- A golden row keeps its evidence: the words that answer it.
- The collection pages count every instance's chunks, not one.
- It runs behind a gateway.
VECTRIXDB_TRUSTED_PROXIESsays whoseX-Forwarded-Forto believe.- The OpenAPI document ships with the release.
- Running heads and feet are taken out, and named.
- A chunk knows whether its own page was read by OCR.
- An HTML table in an extractor's reply becomes rows
compare_extractors- A document fetched from a bucket keeps its pictures.
- A walkthrough that deploys the whole thing on Azure
- Where to stop answering is measured, not borrowed.
evaluation.sweep()compares how documents are cut.vectrixdb golden label- Two homes.
- A figure found by what it looks like.
- The rest of the dashboard's trends
- The Evaluate page has five bands: first, top 3, top 5, top 10, missed.
- A reference extraction server
- A queue between the storage event and the worker on Azure
- The server records how long a search took, and the Overview draws it.
- The Overview and Audit pages have charts.
- Every refusal is one shape.
- An app can act as the person using it:
VECTRIXDB_OIDC_API_AUDIENCE. - A rate limit several servers share, and one for keys.
vectrixdb keys add | list | revoke.vectrixdb check --url- A key can be narrower than a role:
collectionsandexpires_in_days. - A key may arrive as
Authorization: Bearer. - Build an app on it
- Put it behind a gateway
- Masking of email addresses, phone numbers and card numbers.
- Reference pages read off the code.
- Two pages for knowing it all.
- A
testextra - A collection's health says whether it has a knowledge graph
- Evaluate every setup, and pick one of three.
- Golden files from anywhere, and a template to fill.
- A golden file is checked before a run, against one published schema.
- A golden answer can be one page of a document.
- The walkthrough's files are a mirror of the storage account.
vectrixdb golden templateandvectrixdb evaluate.- The Evaluate pages.
- A Function App that evaluates when the golden file changes.
- Azure's semantic ranker, chosen search by search.
- Expected documents are looked up before an evaluation searches.
- Older runs, one menu away.
- A note when a run is missing documents.
- The golden file a run used, for an admin to download.
vectrixdb check, an env file and a template.- The access log on the server's output.
relevance: how good a match is, from 0 to 1, the same on every engine.- The same number from Azure AI Search and OpenSearch.
- A collection's card says what it can do and what it is built with.
- A chunk's text is sent when somebody asks for that chunk.
- Opening a whole document is its own permission,
document.read. - Search excerpts, cut on the server.
- The quality report says why.
- People sign in to the server and its dashboard.
- Passkeys.
- How you sign in.
- Confirm it's you.
- Passwords, only if you want them.
- A lock that grows.
- Single sign-on by security group and an address list.
- Guests and shared collections.
- API keys for scripts.
- Sign-in kept where several servers can share it.
- Passkeys only, if you want that.
- A certificate in place of a client secret.
- A company's own name, logo and colour.
- The dashboard is light by default,
- Roles.
- A signed-in person can search a collection that carries a policy.
- An access log.
- Policy and filter pushdown on OpenSearch.
embeddings=on OpenSearch, with Bedrock.reranker=onVectrix.- A burst of events.
- `vectrixdb.eval...
reranker_en model
ONNX INT8 reranker_en for vectrixdb download-models
dense_en model
ONNX INT8 dense_en for vectrixdb download-models
colbert model
ONNX INT8 colbert for vectrixdb download-models
v2.1.7 - OpenSearch Serverless Compatibility
OpenSearch Serverless Compatibility Fixes
This release fixes critical compatibility issues with AWS OpenSearch Serverless:
Bug Fixes
- Silent error handling - Exceptions from storage backend were being swallowed; now properly re-raised
- Key mapping - Fixed
_embeddingnot mapping todense_embeddingininsert_batch - Custom document IDs - Serverless doesn't support custom
_id; now usingidfield in document body - Refresh policy - Removed
refresh=Truefrom all operations (not supported by Serverless) - Document lookup -
get,update,deletemethods now search byidfield instead of_id
Upgrade Notes
If you're using OpenSearch Serverless, upgrade to this version for proper functionality.
v2.1.5
What's Changed
Proper Hybrid Search with Reciprocal Rank Fusion (RRF)
Implemented industry-standard hybrid retrieval using RRF algorithm:
How it works:
- Stage 1: Run k-NN search (semantic similarity) and BM25 search (lexical matching) independently
- Stage 2: Fuse results using RRF:
score = dense_weight/(k+rank_dense) + sparse_weight/(k+rank_sparse) - Stage 3: Return top results sorted by fused score
Why RRF?
- Robust to score distribution differences between retrieval methods
- No need for score normalization (uses ranks, not raw scores)
- Industry standard (used by Elasticsearch, Cohere, etc.)
- Works with OpenSearch Serverless (no scripting required)
New methods:
text_search()- BM25 lexical search on text_contenthybrid_search()- RRF-fused k-NN + BM25
Parameters:
rrf_k=60- Standard RRF constant from literaturedense_weight=0.7- Weight for semantic resultssparse_weight=0.3- Weight for lexical results
Full Changelog: v2.1.4...v2.1.5
v2.1.3
What's Changed
Fix Hybrid Search Integration
- Fix
_hybrid_searchin Vectrix wrapper to properly detect OpenSearchStorage - Call
storage.hybrid_search()with correct parameters (query_vector,query_text) - Properly handle result format conversion from storage to Vectrix format
Full Changelog: v2.1.2...v2.1.3
v2.1.2
What's Changed
Hybrid Search for OpenSearch
- Add
hybrid_search()method combining k-NN dense vectors with BM25 text matching - Update
vector_search()with optional hybrid mode - Uses OpenSearch script_score to combine semantic + lexical matching
- Configurable weights: dense (default 0.7) and sparse/BM25 (default 0.3)
This enables true hybrid search on OpenSearch Serverless - combining the best of semantic similarity and keyword matching.
Full Changelog: v2.1.1...v2.1.2
v2.1.1
What's Changed
- Fix OpenSearchStorage missing abstract methods (flush, get_collection_config, scan)
- This fixes the TypeError when using VectrixDB with AWS OpenSearch Serverless
Full Changelog: v2.1.0...v2.1.1
v2.1.0: AWS OpenSearch and Aurora PostgreSQL
What's New
AWS Storage Backends
-
OpenSearch Storage - AWS OpenSearch Serverless with native k-NN vector search
- Supports
denseandhybridmodes - Factory method:
VectrixDB.with_opensearch()
- Supports
-
Aurora PostgreSQL Storage - AWS Aurora with pgvector extension
- Supports ALL modes including
ultimate(ColBERT) andgraph - Factory method:
VectrixDB.with_aurora_postgresql()
- Supports ALL modes including
Mode Validation
- OpenSearch now validates mode compatibility and raises a friendly error if you try to use
ultimateorgraphmodes
Installation
pip install vectrixdb[aws] # For AWS backendsQuick Start
from vectrixdb import VectrixDB, Vectrix
# OpenSearch (dense/hybrid only)
opensearch = VectrixDB.with_opensearch(
endpoint="https://xxx.us-east-1.aoss.amazonaws.com",
region="us-east-1",
)
db = Vectrix("products", mode="hybrid", storage_backend=opensearch)
# Aurora PostgreSQL (all modes)
aurora = VectrixDB.with_aurora_postgresql(
host="cluster.xxx.rds.amazonaws.com",
user="admin",
password="password",
)
db = Vectrix("products", mode="ultimate", storage_backend=aurora)