Skip to content

Getting Started

Doug Gerard edited this page May 14, 2026 · 1 revision

Getting Started

This guide takes you from zero to your first successful documentation search in Claude Code.


Prerequisites

Before installing SaddleRAG, you need:

Requirement Notes
Windows 10 version 1903 or later The installer enforces this minimum; DirectML GPU acceleration requires it
MongoDB Community Edition 6.0+ works fine. The installer does not install MongoDB — you must have it running before installing SaddleRAG. Default: mongodb://localhost:27017
Ollama Required for page classification. Install from ollama.com. After installing, run ollama pull phi4-mini to pre-download the default classification model.
Claude Code, VS Code with Copilot, or Claude Desktop SaddleRAG integrates with these tools; at least one must be installed

GPU note: The installer detects your GPU adapter and pre-selects DirectML (GPU-accelerated) if a capable GPU is found. GPU acceleration speeds up embedding at ingestion time. It is optional — CPU mode works correctly but is slower for large scrape jobs.


Installation

From the MSI installer (recommended)

  1. Download the latest SaddleRAG.Mcp-{version}.msi from the Releases page
  2. Run the installer
  3. On the MongoDB page, enter your connection string and click Test Connection — proceed only after seeing a green checkmark
  4. On the Ollama page, verify Ollama is running and click Test Connection
  5. Complete the installation — the installer will:
    • Install SaddleRAG to %ProgramFiles%\SaddleRAG\
    • Create and start a Windows service named SaddleRAGMcp
    • Download the ONNX embedding model (~273 MB) and reranker model (~244 MB) to %ProgramData%\SaddleRAG\models\onnx\
    • Register SaddleRAG in Claude Code, Claude Desktop, VS Code, and GitHub Copilot CLI

Verify the installation

Open a browser and navigate to http://localhost:6100/monitor. You should see the SaddleRAG monitor dashboard. The warmup status should transition from Starting → DownloadingModels → LoadingVectorIndex → Ready over a minute or two (longer on first install while ONNX models download).

Alternatively, run:

curl http://localhost:6100/health

A "Status":"Healthy" response confirms the service is running.

Build from source

git clone https://github.com/JackalopeTechnologies/SaddleRAG.git
cd SaddleRAG

# GPU-enabled build (DirectML)
dotnet build SaddleRAG.slnx

# CPU-only build
dotnet build SaddleRAG.slnx -p:UseGpu=false

# Run
dotnet run --project SaddleRAG.Mcp

First use: index a library

Via an AI assistant (recommended)

Open Claude Code and ask:

"Index the documentation for Polly, the .NET resilience library. The docs are at https://www.pollydocs.org/"

SaddleRAG registers a start_ingest tool that Claude Code will call. The tool is a state machine that guides Claude through any needed steps (reconnaissance, URL confirmation) before firing the scrape. You'll see progress in Claude's tool call output. The scrape runs in the background — you can continue working while it proceeds.

To check progress, ask:

"What's the status of the Polly scrape?"

Or visit http://localhost:6100/monitor.

Via the CLI

# Index with explicit URL and library details
saddlerag ingest --url https://www.pollydocs.org/ --library polly --version 8.2.0

# Check scrape status
saddlerag status

# Dry run: preview what would be crawled without writing to the database
saddlerag dryrun --url https://www.pollydocs.org/ --library polly --version 8.2.0

First use: search

Once a scrape completes, ask your AI assistant questions that require the documentation:

"How do I configure a retry strategy with exponential backoff in Polly 8?"

SaddleRAG's search_docs tool will be called automatically. The assistant receives the relevant documentation chunks and can answer accurately.

You can also search directly:

"Search the Polly docs for AddRetry"

Or use the search_docs tool explicitly with the AI:

"Use search_docs to find examples of ResiliencePipeline in the Polly docs"


Registering AI client integrations manually

If you built from source or need to re-register the AI tool integrations:

# Register with all supported AI tools
saddlerag register-clients

# Unregister (remove SaddleRAG from all AI tool configs)
saddlerag unregister-clients

register-clients writes MCP entries to:

  • ~/.claude.json (Claude Code)
  • %APPDATA%\Claude\claude_desktop_config.json (Claude Desktop)
  • VS Code user settings.json
  • GitHub Copilot CLI config

After registering, restart your AI tool for the changes to take effect.


Next steps


Troubleshooting

"Not Found" when calling MCP tools

SaddleRAG is not running or not reachable at port 6100. Check:

Get-Service SaddleRAGMcp
# If Stopped:
Start-Service SaddleRAGMcp

Classification is slow

Ollama is downloading the phi4-mini model on first use (~2.4 GB). Wait for the download to complete, then re-run the scrape. Check Ollama status at http://localhost:11434.

"WARMING_UP" response from search tools

The server is still loading the vector index from MongoDB. Wait for the health endpoint to return "WarmupStatus":"Completed".

MongoDB connection failure

Verify MongoDB is running:

Get-Service MongoDB

Check the connection string in %ProgramFiles%\SaddleRAG\SaddleRAG.Mcp\appsettings.json.

Clone this wiki locally