-
Notifications
You must be signed in to change notification settings - Fork 0
Hybrid Search
SaddleRAG uses a three-layer retrieval pipeline: vector (semantic) search, BM25 (keyword) search, and optional cross-encoder reranking. Each layer compensates for the weaknesses of the others.
Neither pure vector search nor pure keyword search is ideal for documentation retrieval.
Vector search alone is good at semantic similarity. "Configuring retry behavior" will find content about "retry policies" even without exact word overlap. But it can miss content when the query uses an exact method name or class name that appears verbatim in the docs. If you search for ResiliencePipeline.AddRetry, a pure vector search may rank "retry patterns overview" above the actual AddRetry API reference.
BM25 (keyword) search alone is good at exact term matching. ResiliencePipeline.AddRetry will surface docs that contain exactly those tokens. But it misses semantic similarity — a query about "exponential backoff" won't find content that uses the word "delay-multiplier" without also mentioning "backoff."
Hybrid search blends the two scores. The result is better than either alone: you get semantic coverage for natural-language questions and exact-match precision for identifier-style queries.
Reranking adds a third pass for the top candidates, using a cross-encoder model that considers query and document together. This corrects cases where the hybrid score ranked a less-relevant result above a more-relevant one.
The query string is embedded using the active embedding provider (ONNX by default). For ONNX with nomic-embed-text-v1.5, the query gets the asymmetric task prefix:
search_query: how do I configure retry policies in Polly 8?
The result is a float[768] vector.
The query vector is compared against every chunk in the in-memory index for the requested (library, version) — or across all libraries if no filter is specified. Similarity is measured by cosine similarity:
cosine_similarity(query_vector, chunk_vector) = dot(q, c)
(Both vectors are L2-normalized, so the dot product equals cosine similarity.)
The number of candidates retrieved is max(maxResults, 25) × 5. If you request 10 results, the vector search retrieves 125 candidates. This over-retrieval gives the hybrid blend and reranking steps enough candidates to work with.
Results are sorted by vector score descending.
BM25 (Best Match 25) is a classical information retrieval ranking function that scores documents based on term frequency, inverse document frequency, and document length normalization.
The query is tokenized two ways:
- Word tokens — lowercased words split on whitespace and punctuation
- Identifier tokens — sequences that look like code identifiers (alphanumeric + underscore, mixed case)
For each token, the BM25 scorer looks up a postings list in the sharded BM25 index. The BM25 score for each chunk that contains at least one query term is computed using the standard BM25 formula with the corpus statistics stored in LibraryIndex.
BM25 scoring only runs for single-library, single-version queries. Cross-library searches (no library filter) use vector search only.
Vector and BM25 scores are combined into a single hybrid score. BM25 scores are first normalized to [0, 1] by dividing by the maximum BM25 score in this result set. Then:
hybridScore = (1 - BM25Weight) × vectorScore + BM25Weight × normalizedBM25Score
The default BM25Weight is 0.4, giving 60% weight to vector similarity and 40% to keyword relevance. This balance was chosen because:
- Documentation queries are often semantic ("how do I do X") but frequently contain specific API identifiers
- 60/40 gives strong semantic coverage without losing the precision boost from keyword matches
The weight is configurable via Ranking.Bm25Weight in appsettings.json. A weight of 0.0 gives pure vector search; 1.0 gives pure BM25.
Results are sorted by hybrid score descending and truncated to maxResults × 5 to pass to the reranker.
Enabled when: Ranking.ReRankerStrategy = Onnx and there are at least 6 hybrid candidates.
The cross-encoder reranker takes (query, document) pairs and produces a relevance score for each. Unlike the bi-encoder embedding model (which embeds query and document independently and measures distance), the cross-encoder processes the full sequence [CLS] query [SEP] document [SEP] in one forward pass. This lets the model directly attend to the relationship between the query and each document.
The tradeoff: cross-encoders are ~10× slower than bi-encoders for the same text length. That's why reranking runs only on the top candidates (default 12), not the full result set.
The reranking process:
- Take the top
MaxReRankCandidates(default 12) hybrid results - For each candidate, assemble the pair string:
[CLS] {query} [SEP] {chunk_content} [SEP] - Tokenize with SentencePiece (the mxbai-rerank-base-v1 tokenizer), truncating to 512 tokens
- Batch the pairs (default batch size 64) and run them through the ONNX cross-encoder
- The output is a scalar logit for each pair; apply sigmoid to map to (0, 1)
- Sort the top-12 candidates by sigmoid score (descending)
- Append any hybrid candidates beyond the top-12 (in their hybrid order) after the reranked tier
- Truncate to
maxResults
When to enable reranking:
Reranking adds ~100–300ms to search latency (depending on hardware and candidate count). For interactive use where the AI assistant is waiting, this is usually acceptable. Enable it if you want the highest quality results and can tolerate the extra latency.
To enable: set_rerank_strategy MCP tool, or set Ranking.ReRankerStrategy = Onnx in appsettings.json. This takes effect at runtime without a server restart.
Every search_docs response includes a Strategy block and a Timing block that explain exactly what happened:
{
"Strategy": {
"ReRankerStrategy": "Onnx",
"RerankActive": true,
"QueryIsIdentifierShape": false,
"Category": null,
"Bm25Weight": 0.4
},
"Timing": {
"EmbedMs": 12,
"VectorSearchMs": 45,
"Bm25Ms": 8,
"ReRankMs": 187,
"TotalMs": 252
}
}RerankActive can be false even when strategy is Onnx if there weren't enough candidates. QueryIsIdentifierShape is true when the query looks like an identifier (e.g., ResiliencePipeline.AddRetry) — this enables a symbol backstop path that searches by identifier name directly, in addition to vector search.
When QueryIsIdentifierShape is true, SaddleRAG runs an additional retrieval path alongside the vector and BM25 searches: it queries the chunks collection for chunks whose QualifiedName or Symbols list contains the query text exactly. This is a guaranteed hit for any API type or method that was indexed — even if the vector and BM25 searches ranked it lower than expected.
Symbol backstop results are merged into the candidate set before hybrid blending.
The search_docs tool accepts an optional category parameter. When set, only chunks of that category are returned. This is useful for:
-
category: ApiReference— find exact method/class documentation -
category: Sample— find working code examples only -
category: HowTo— find step-by-step guides
Category filtering happens during the vector search step, reducing the candidate set early. BM25 and reranking operate only on the filtered candidates.
All search configuration lives in appsettings.json under Ranking:
| Setting | Default | Effect |
|---|---|---|
Bm25Weight |
0.4 |
BM25 fraction in hybrid blend. 0.0 = pure vector, 1.0 = pure BM25. |
VectorCandidateMultiplier |
5 |
Vector candidates = max(requested, MinVec) × multiplier
|
MinVectorCandidateCount |
25 |
Floor on vector candidates regardless of requested result count |
MaxReRankCandidates |
12 |
Top-N candidates passed to cross-encoder |
ReRankerStrategy |
Off |
Off or Onnx
|
ReRankerStrategy can be changed at runtime without restart via the set_rerank_strategy MCP tool.