Skip to content

Maintenance and Operations

Doug Gerard edited this page May 14, 2026 · 1 revision

Maintenance and Operations

This page covers day-to-day operations: keeping documentation indexes current, fixing quality issues, monitoring health, managing versions, and running SaddleRAG as a Windows service.


Keeping documentation current

Documentation changes as libraries release new versions. SaddleRAG doesn't auto-update indexes — you control when re-indexing happens.

Refresh an existing index

When a library publishes a new version:

Via AI assistant:

"Rescrape the Polly docs for version 8.3.0"

The assistant will call start_ingest with the new version, which may trigger recon_library if this is a version that differs significantly from the previously indexed one, then scrape_docs.

Via CLI:

saddlerag ingest --url https://www.pollydocs.org/ --library polly --version 8.3.0

Rescrape from stored URLs

If you want to refresh an index without changing the version (e.g., the docs were corrected without a version bump):

rescrape_library(library="polly", version="8.2.0")

This re-fetches all stored page URLs for the version without requiring you to re-enter the root URL and crawl options. New and changed content is re-classified, re-chunked, and re-embedded. Pages whose content hasn't changed since the last scrape are skipped (content hash comparison).

Version management

SaddleRAG indexes each version independently. A library can have multiple versions indexed simultaneously — list_libraries shows all available versions. AI assistants default to the CurrentVersion when no version is specified.

To compare what changed between versions:

get_version_changes(library="polly", fromVersion="8.1.0", toVersion="8.2.0")

Returns lists of added pages, removed pages, and changed pages (by content hash), with a brief summary of changes. This is useful for understanding what documentation work a library version update brings.

To delete an old version that's no longer needed:

delete_version(library="polly", version="7.0.0", dryRun=false)

Fixing quality issues

Diagnosing poor search results

Start with get_library_health:

get_library_health(library="polly", version="8.2.0")

Look for:

  • High Unclassified fraction — classification may have failed during the scrape (was Ollama running? was the model available?)
  • Very low chunk count relative to page count — pages may have been fetched as empty (JavaScript rendering failed, or the site requires authentication)
  • Boundary issue rate > 10% — chunks are being split mid-sentence more than expected; may indicate the chunking strategy isn't matching the content structure
  • Suspect markers — flags from SuspectDetector indicating implausible scrape results

Resolving classification failures

If many pages are Unclassified because Ollama was unavailable during the scrape:

  1. Ensure Ollama is running with the classification model loaded
  2. Run rechunk_library with reclassify: true:
    rechunk_library(library="polly", version="8.2.0", reclassify=true)
    
    This re-runs classification on all stored pages, then re-chunks with the correct strategy, then re-embeds. It does not re-crawl.

Resolving empty pages

If chunks have very little content (pages fetched as empty):

  1. Check list_pages(library="polly", version="8.2.0") — look for pages with 0 or 1 chunks
  2. Check whether the documentation site uses JavaScript rendering; if SaddleRAG's Playwright browser isn't executing the JS, pages will appear empty
  3. Try add_page(url="...") for specific problem pages to see what content is fetched
  4. If many pages are empty, the site may require authentication; URL patterns may need adjustment to exclude authenticated areas

Re-indexing with a new embedding model

If you want to switch to a better embedding model:

  1. Configure the new model (see ONNX Models and GPU)
  2. Restart SaddleRAGMcp to load the new model
  3. Re-embed all libraries:
    reembed_library(library="polly", version="8.2.0")
    
    Do this for each (library, version) you have indexed.

After re-embedding, the in-memory vector index is rebuilt with the new vectors. Search results will reflect the new model.

Updating the library profile

If a library's LibraryProfile was wrong (missing symbols, incorrect URL patterns):

  1. Generate a corrected profile using recon_library or submit_library_profile
  2. Run reextract_library to re-run symbol extraction with the updated profile
  3. If URL patterns changed, run rescrape_library to re-crawl with the corrected patterns

Monitoring

The Blazor monitor UI

Navigate to http://localhost:6100/monitor in a browser. The monitor shows:

  • Active scrape jobs with real-time progress (pages crawled, classified, chunked, embedded)
  • Recent job history with final status
  • Per-stage timing for completed jobs
  • Server warmup status and phases

The monitor updates in real time via SignalR while a job is running.

Health endpoint

GET http://localhost:6100/health

Returns JSON:

{
  "Status": "Healthy",
  "WarmupStatus": "Completed",
  "WarmupPhase": null,
  "WarmupError": null
}

Use this endpoint in monitoring infrastructure to alert if SaddleRAG goes unhealthy.

Logs

Log files: %LocalAppData%\SaddleRAG\logs\saddlerag-{date}.log

Retained for 7 days, rolling daily. Level defaults to Information. Adjust via toggle_logging MCP tool or Logging.LogLevel in appsettings.json.

When running as a Windows service, events at Warning and above are also written to the Windows Event Log under Application.

Dashboard overview

Call get_dashboard_index from any AI assistant for a quick health check:

get_dashboard_index()

The SuggestedNextAction field tells you if anything needs attention (stale running job, suspect library, model download needed, etc.).


Windows service operations

SaddleRAG installs as a Windows service named SaddleRAGMcp.

# Check service status
Get-Service SaddleRAGMcp

# Start / stop / restart
Start-Service SaddleRAGMcp
Stop-Service SaddleRAGMcp
Restart-Service SaddleRAGMcp

# View service event log entries
Get-EventLog -LogName Application -Source SaddleRAGMcp -Newest 50

Service startup type

The service installs as Automatic (starts with Windows). To change:

Set-Service SaddleRAGMcp -StartupType Manual

Service identity

The service runs as Local System by default (installer-configured). For a team server with MongoDB authentication, you may want to run it under a domain service account that has appropriate MongoDB credentials.


Database maintenance

Audit log cleanup

The scrape_audit_log collection has a 30-day TTL index — entries auto-expire. To force early cleanup:

cleanup_audit_log()     # MCP tool

Orphan chunk cleanup

If scrape jobs were interrupted repeatedly, there may be chunks in MongoDB without a corresponding LibraryVersionRecord. Run:

cleanup_orphans()       # MCP tool

This identifies chunks whose (library, version) has no LibraryVersionRecord and removes them.

Job history cleanup

Scrape jobs and background jobs also expire after 30 days via TTL. cleanup_jobs manually removes completed and failed jobs older than the configured retention period.


Backup and recovery

SaddleRAG's data lives entirely in MongoDB. To back up:

mongodump --db SaddleRAG --out "E:\backups\saddlerag-$(Get-Date -Format 'yyyyMMdd')"

To restore:

mongorestore --db SaddleRAG "E:\backups\saddlerag-20260514"

After restoring, restart SaddleRAGMcp so it reloads the vector index from the restored data.

The ONNX model files (%ProgramData%\SaddleRAG\models\onnx\) do not need to be backed up — they are downloaded from Hugging Face automatically on startup.


Upgrading SaddleRAG

Run the new MSI installer over the existing installation. The installer:

  1. Stops the SaddleRAGMcp service
  2. Replaces the binary files
  3. Preserves appsettings.json and runtime-overrides.json
  4. Re-downloads ONNX models if the configured models changed
  5. Restarts the service

After upgrading, check get_dashboard_index to verify the new version is running and healthy. If the upgrade changes the chunking algorithm or embedding model version, you may want to run rechunk_library or reembed_library for your indexed libraries to take advantage of improvements.

Clone this wiki locally