-
Notifications
You must be signed in to change notification settings - Fork 0
Getting Started
This guide takes you from zero to your first successful documentation search in Claude Code.
Before installing SaddleRAG, you need:
| Requirement | Notes |
|---|---|
| Windows 10 version 1903 or later | The installer enforces this minimum; DirectML GPU acceleration requires it |
| MongoDB | Community Edition 6.0+ works fine. The installer does not install MongoDB — you must have it running before installing SaddleRAG. Default: mongodb://localhost:27017
|
| Ollama | Required for page classification. Install from ollama.com. After installing, run ollama pull phi4-mini to pre-download the default classification model. |
| Claude Code, VS Code with Copilot, or Claude Desktop | SaddleRAG integrates with these tools; at least one must be installed |
GPU note: The installer detects your GPU adapter and pre-selects DirectML (GPU-accelerated) if a capable GPU is found. GPU acceleration speeds up embedding at ingestion time. It is optional — CPU mode works correctly but is slower for large scrape jobs.
- Download the latest
SaddleRAG.Mcp-{version}.msifrom the Releases page - Run the installer
- On the MongoDB page, enter your connection string and click Test Connection — proceed only after seeing a green checkmark
- On the Ollama page, verify Ollama is running and click Test Connection
- Complete the installation — the installer will:
- Install SaddleRAG to
%ProgramFiles%\SaddleRAG\ - Create and start a Windows service named
SaddleRAGMcp - Download the ONNX embedding model (~273 MB) and reranker model (~244 MB) to
%ProgramData%\SaddleRAG\models\onnx\ - Register SaddleRAG in Claude Code, Claude Desktop, VS Code, and GitHub Copilot CLI
- Install SaddleRAG to
Open a browser and navigate to http://localhost:6100/monitor. You should see the SaddleRAG monitor dashboard. The warmup status should transition from Starting → DownloadingModels → LoadingVectorIndex → Ready over a minute or two (longer on first install while ONNX models download).
Alternatively, run:
curl http://localhost:6100/health
A "Status":"Healthy" response confirms the service is running.
git clone https://github.com/JackalopeTechnologies/SaddleRAG.git
cd SaddleRAG
# GPU-enabled build (DirectML)
dotnet build SaddleRAG.slnx
# CPU-only build
dotnet build SaddleRAG.slnx -p:UseGpu=false
# Run
dotnet run --project SaddleRAG.McpOpen Claude Code and ask:
"Index the documentation for Polly, the .NET resilience library. The docs are at https://www.pollydocs.org/"
SaddleRAG registers a start_ingest tool that Claude Code will call. The tool is a state machine that guides Claude through any needed steps (reconnaissance, URL confirmation) before firing the scrape. You'll see progress in Claude's tool call output. The scrape runs in the background — you can continue working while it proceeds.
To check progress, ask:
"What's the status of the Polly scrape?"
Or visit http://localhost:6100/monitor.
# Index with explicit URL and library details
saddlerag ingest --url https://www.pollydocs.org/ --library polly --version 8.2.0
# Check scrape status
saddlerag status
# Dry run: preview what would be crawled without writing to the database
saddlerag dryrun --url https://www.pollydocs.org/ --library polly --version 8.2.0Once a scrape completes, ask your AI assistant questions that require the documentation:
"How do I configure a retry strategy with exponential backoff in Polly 8?"
SaddleRAG's search_docs tool will be called automatically. The assistant receives the relevant documentation chunks and can answer accurately.
You can also search directly:
"Search the Polly docs for AddRetry"
Or use the search_docs tool explicitly with the AI:
"Use search_docs to find examples of ResiliencePipeline in the Polly docs"
If you built from source or need to re-register the AI tool integrations:
# Register with all supported AI tools
saddlerag register-clients
# Unregister (remove SaddleRAG from all AI tool configs)
saddlerag unregister-clientsregister-clients writes MCP entries to:
-
~/.claude.json(Claude Code) -
%APPDATA%\Claude\claude_desktop_config.json(Claude Desktop) - VS Code user
settings.json - GitHub Copilot CLI config
After registering, restart your AI tool for the changes to take effect.
-
Index more libraries — use
scrape_docsfor each library your project uses - Understand what happened — how the five-stage pipeline works
- Enable GPU acceleration — faster ingestion with DirectML or CUDA
- Set up team sharing — share a SaddleRAG instance with colleagues
- Tune search quality — enable cross-encoder reranking for better results
SaddleRAG is not running or not reachable at port 6100. Check:
Get-Service SaddleRAGMcp
# If Stopped:
Start-Service SaddleRAGMcpOllama is downloading the phi4-mini model on first use (~2.4 GB). Wait for the download to complete, then re-run the scrape. Check Ollama status at http://localhost:11434.
The server is still loading the vector index from MongoDB. Wait for the health endpoint to return "WarmupStatus":"Completed".
Verify MongoDB is running:
Get-Service MongoDBCheck the connection string in %ProgramFiles%\SaddleRAG\SaddleRAG.Mcp\appsettings.json.