Skip to content
Doug Gerard edited this page May 14, 2026 · 1 revision

SaddleRAG

SaddleRAG is a local documentation indexing and retrieval server that gives AI coding assistants accurate, up-to-date knowledge of the libraries and APIs you actually use — even when those libraries are niche, internal, or newer than the assistant's training data.

It runs on your machine (or your team's server), crawls documentation sites, and exposes everything through the Model Context Protocol (MCP) so that Claude Code, Claude Desktop, VS Code Copilot, and GitHub Copilot CLI can query it in real time.


The one-line version

Your AI assistant knows general coding patterns but not your specific library versions. SaddleRAG fixes that.


What it does

  1. Crawls any documentation website with a headless browser, following links breadth-first
  2. Classifies each page (Overview, HowTo, API reference, code sample, etc.) using a local LLM
  3. Chunks the content using a strategy tuned to each page type
  4. Embeds every chunk into a 768-dimensional vector using an in-process ONNX model
  5. Indexes everything in MongoDB alongside a BM25 full-text index
  6. Serves search results over MCP so your AI assistant can query documentation in real time, mid-conversation

The result: when you ask your AI assistant "how do I configure retry policies in Polly 8?" it looks up the answer in your locally-indexed docs rather than guessing from training data.


Key design choices

What Why
Ollama for page classification Classification drives chunking strategy — getting it right is more important than getting it fast. A small local LLM (phi4-mini) does this accurately without cloud costs or data egress.
ONNX for embeddings, not Ollama In-process ONNX inference has zero IPC overhead and supports asymmetric bi-encoder prompting (different prefixes for indexing vs. querying), which materially improves retrieval accuracy.
ONNX cross-encoder for reranking Cross-encoder reranking is a different architecture than bi-encoder embedding — no Ollama model covers it. The ONNX cross-encoder rescore the top candidates with query-aware pair scoring.
In-memory vector search Documentation corpora fit in RAM. No MongoDB Atlas, no vector database subscription needed.
Hybrid vector + BM25 Vector search misses exact keyword matches; BM25 misses semantic similarity. Blending both gives consistently better results than either alone.
Multi-profile MongoDB One MCP server can serve multiple teams or environments simultaneously via named connection profiles.

Wiki contents

Understanding SaddleRAG

Using SaddleRAG

How it works (in depth)

Operating SaddleRAG


Licensing

SaddleRAG is source-available. It is free for a single individual developer. Team and commercial use (multiple authorized users sharing one instance) requires a commercial license at $100 per authorized user per year.

Clone this wiki locally