-
Notifications
You must be signed in to change notification settings - Fork 0
Home
SaddleRAG is a local documentation indexing and retrieval server that gives AI coding assistants accurate, up-to-date knowledge of the libraries and APIs you actually use — even when those libraries are niche, internal, or newer than the assistant's training data.
It runs on your machine (or your team's server), crawls documentation sites, and exposes everything through the Model Context Protocol (MCP) so that Claude Code, Claude Desktop, VS Code Copilot, and GitHub Copilot CLI can query it in real time.
Your AI assistant knows general coding patterns but not your specific library versions. SaddleRAG fixes that.
- Crawls any documentation website with a headless browser, following links breadth-first
- Classifies each page (Overview, HowTo, API reference, code sample, etc.) using a local LLM
- Chunks the content using a strategy tuned to each page type
- Embeds every chunk into a 768-dimensional vector using an in-process ONNX model
- Indexes everything in MongoDB alongside a BM25 full-text index
- Serves search results over MCP so your AI assistant can query documentation in real time, mid-conversation
The result: when you ask your AI assistant "how do I configure retry policies in Polly 8?" it looks up the answer in your locally-indexed docs rather than guessing from training data.
| What | Why |
|---|---|
| Ollama for page classification | Classification drives chunking strategy — getting it right is more important than getting it fast. A small local LLM (phi4-mini) does this accurately without cloud costs or data egress. |
| ONNX for embeddings, not Ollama | In-process ONNX inference has zero IPC overhead and supports asymmetric bi-encoder prompting (different prefixes for indexing vs. querying), which materially improves retrieval accuracy. |
| ONNX cross-encoder for reranking | Cross-encoder reranking is a different architecture than bi-encoder embedding — no Ollama model covers it. The ONNX cross-encoder rescore the top candidates with query-aware pair scoring. |
| In-memory vector search | Documentation corpora fit in RAM. No MongoDB Atlas, no vector database subscription needed. |
| Hybrid vector + BM25 | Vector search misses exact keyword matches; BM25 misses semantic similarity. Blending both gives consistently better results than either alone. |
| Multi-profile MongoDB | One MCP server can serve multiple teams or environments simultaneously via named connection profiles. |
- Why SaddleRAG — The problem: training cutoffs, niche libraries, context limits
- Theory of Operation — Full architecture and data flow
- Getting Started — Installation, prerequisites, first scrape, first search
- MCP Tools Reference — All tools the MCP server exposes to AI assistants
- CLI Reference — Command-line interface for scripting and automation
- Ingestion Pipeline — The five-stage crawl → classify → chunk → embed → index pipeline
- Page Classification — How Ollama classifies pages and why it matters
- Embeddings: ONNX vs Ollama — The embedding design decision explained
- Hybrid Search — Vector search + BM25 + cross-encoder reranking
- Library Reconnaissance — The recon step that shapes how a library is indexed
- Profiles and Team Usage — Sharing a SaddleRAG instance across a team
- ONNX Models and GPU — Models, execution providers, GPU acceleration
- Maintenance and Operations — Re-indexing, version management, health monitoring
-
Configuration Reference — Every setting in
appsettings.json
SaddleRAG is source-available. It is free for a single individual developer. Team and commercial use (multiple authorized users sharing one instance) requires a commercial license at $100 per authorized user per year.