Skip to content

MCP Tools Reference

Doug Gerard edited this page May 14, 2026 · 1 revision

MCP Tools Reference

SaddleRAG exposes its functionality to AI assistants through the Model Context Protocol (MCP). Tools are served at http://localhost:6100/mcp using the Streamable HTTP transport.

Six tools are marked alwaysLoad and are available in every AI assistant session without the assistant needing to discover them first. The rest are loaded on demand.


Always-loaded tools

These tools are injected into every session automatically.


get_dashboard_index

Purpose: Single-call session startup overview. Call this first in any session where you might need documentation.

Returns:

  • Library and version counts across all profiles
  • Recent scrape jobs with status and progress
  • Server warmup status and health
  • Suspect library flags
  • SuggestedNextAction — the highest-priority tool to call next, if any

This is the right first call in any session. It tells you what's indexed, whether any scrapes are running or failed, and whether any action is needed before querying.

Parameters:

  • profile (optional) — target a specific MongoDB profile

list_libraries

Purpose: Returns all indexed libraries, their current version, and all available versions.

Use this to discover what documentation is available before searching.

Parameters:

  • profile (optional)

search_docs

Purpose: Natural language search across indexed documentation.

The most frequently called tool. Runs hybrid vector + BM25 retrieval with optional cross-encoder reranking.

Parameters:

  • query — the search query
  • library (optional) — restrict to a specific library
  • version (optional) — restrict to a specific version; defaults to CurrentVersion when library is set
  • category (optional) — restrict to a specific category (Overview, HowTo, Sample, Code, ApiReference, ChangeLog)
  • maxResults (optional, default 5) — number of results to return
  • profile (optional)

Returns: Result chunks with relevance scores, page URLs, section paths, and timing/strategy metadata.


get_class_reference

Purpose: Look up an API type by name. Tries exact match first, then fuzzy match.

Optimized for "show me the docs for class X" queries. Searches across all libraries unless filtered.

Parameters:

  • className — the class or type name to look up (partial names work)
  • library (optional) — restrict to a specific library
  • version (optional)
  • profile (optional)

get_library_overview

Purpose: Get conceptual overview content for a library. Returns chunks from the Overview category with fallback to all categories if no Overview content exists.

Parameters:

  • library — the library identifier
  • version (optional)
  • profile (optional)

list_symbols

Purpose: List documented symbols for a library, filtered by kind.

Parameters:

  • library — the library identifier
  • kind (optional) — class, enum, function, or parameter
  • filter (optional) — partial name filter (case-insensitive)
  • version (optional)
  • profile (optional)

Ingestion tools


start_ingest

Purpose: The front door for indexing a new library. This is a state machine — call it first, then follow the NextTool and NextToolArgs it returns.

Possible return states:

  • IN_PROGRESS — a scrape is already running for this library/version
  • URL_SUSPECT — the URL looks wrong (404, login wall, etc.); suggests correction
  • RECON_NEEDED — no LibraryProfile exists; call recon_library next
  • READY_TO_SCRAPE — everything is set up; call scrape_docs next
  • STALE — library exists but docs may be outdated; offers rescrape
  • READY — library is already indexed and up to date

Parameters:

  • url — the root URL of the documentation site
  • library — a short identifier for the library (e.g., polly, serilog)
  • version — the version being indexed (e.g., 8.2.0)
  • profile (optional)

scrape_docs

Purpose: Start a scrape job. Runs in the background.

Parameters:

  • url — root URL
  • library — library identifier
  • version — version string
  • allowedUrlPatterns (optional) — regex patterns; only matching URLs are crawled
  • excludedUrlPatterns (optional) — regex patterns; matching URLs are skipped
  • crawlBudget (optional) — maximum number of pages to fetch
  • profile (optional)

rescrape_library

Purpose: Re-scrape from the stored page URLs for an existing library version. Useful for refreshing an outdated index without having to re-enter the URL and options.

Parameters:

  • library
  • version (optional, defaults to current)
  • profile (optional)

dryrun_scrape

Purpose: Preview the crawl scope without writing anything to the database. Returns the list of URLs that would be fetched and the estimated page count.

Parameters: Same as scrape_docs


index_project_dependencies

Purpose: Scan the current project's dependency files (.csproj, package.json, requirements.txt) and queue scrape jobs for any dependencies that aren't already indexed.

Supports NuGet, npm, and pip ecosystems. SaddleRAG resolves documentation URLs from the registry metadata for each package.

Parameters:

  • projectPath — path to the project root
  • profile (optional)

Job management tools


get_scrape_status

Purpose: Get the status and progress of a running or recent scrape job.

Parameters:

  • jobId — from the scrape_docs response, or use list_scrape_jobs to find it
  • profile (optional)

list_scrape_jobs

Purpose: List all recent scrape jobs with their status.

Parameters:

  • profile (optional)
  • limit (optional)

cancel_scrape

Purpose: Cancel a running scrape job.

Parameters:

  • jobId
  • profile (optional)

Library and page tools


list_pages

Purpose: List the pages stored for a library version with their URL, title, category, and chunk count.

Parameters:

  • library
  • version (optional)
  • category (optional) — filter by category
  • profile (optional)

add_page

Purpose: Add a single page to an existing library index without running a full scrape. Useful for adding a specific URL that the crawler missed or for one-off additions.

Parameters:

  • url — the URL to fetch and index
  • library
  • version (optional)
  • profile (optional)

get_version_changes

Purpose: Compare two versions of a library and return a diff of what pages were added, removed, or changed.

Parameters:

  • library
  • fromVersion
  • toVersion
  • profile (optional)

Health tools


get_library_health

Purpose: Diagnostic report for a library's indexed content. Returns:

  • Total chunk count and page count
  • Chunk count by category
  • Language distribution (programming languages detected in code chunks)
  • Hostname distribution (fraction of pages from each domain)
  • Boundary issue rate (chunks that appear to be split mid-sentence)
  • Suspect markers (flags from SuspectDetector)

Use this when search quality seems poor for a library — health issues often explain the problem.

Parameters:

  • library
  • version (optional)
  • profile (optional)

Library administration tools

These tools default to dryRun: true. Set dryRun: false to actually execute.


rename_library

Purpose: Rename a library's identifier across all collections. All pages, chunks, and metadata are updated atomically.

Parameters:

  • oldLibraryId
  • newLibraryId
  • dryRun (default true)
  • profile (optional)

delete_version

Purpose: Delete a specific version of a library. Removes all pages, chunks, BM25 shards, and metadata for that version.

Parameters:

  • library
  • version
  • dryRun (default true)
  • profile (optional)

delete_library

Purpose: Delete a library entirely — all versions, all data.

Parameters:

  • library
  • dryRun (default true)
  • profile (optional)

Index maintenance tools


rechunk_library

Purpose: Re-chunk all stored pages without re-crawling. Use this after:

  • The chunking algorithm was improved
  • Classification results were wrong and have been corrected
  • The library profile changed (new symbols, different casing conventions)

Parameters:

  • library
  • version (optional)
  • reclassify (optional, default false) — re-run Ollama classification before rechunking
  • profile (optional)

reembed_library

Purpose: Re-embed all stored chunks using the currently active embedding provider. Use this after:

  • Switching embedding providers (ONNX → Ollama or vice versa)
  • Switching to a different embedding model
  • The embedding model was updated and vectors may have shifted

Parameters:

  • library
  • version (optional)
  • profile (optional)

reextract_library

Purpose: Re-run symbol extraction on existing chunks without re-chunking or re-embedding. Use this after updating the library profile's symbol list or stoplist.

Parameters:

  • library
  • version (optional)
  • profile (optional)

recon_library

Purpose: Run library reconnaissance. This analyzes the library's documentation structure and produces a LibraryProfile that improves chunking and symbol extraction.

Normally called by the AI assistant as part of the start_ingest state machine. The AI assistant fills in the LibraryProfile fields based on its knowledge of the library; this tool persists the profile.

Parameters: See Library Reconnaissance for the full schema.


submit_library_profile

Purpose: Directly submit a completed LibraryProfile for a library. Used when the AI assistant has generated the profile and is ready to persist it.


Symbol management tools


list_excluded_symbols

Purpose: List symbols that have been manually added to the stoplist or likely-symbols list for a library.

Parameters:

  • library
  • version (optional)
  • profile (optional)

add_to_likely_symbols

Purpose: Add a symbol to the library's LikelySymbols list. Symbols in this list are actively sought by the symbol extractor and given priority in search.

Parameters:

  • library, version, symbolName, profile (optional)

add_to_stoplist

Purpose: Add a term to the library's stoplist. Stoplist terms are ignored by the symbol extractor and BM25 indexer — useful for common words that appear in this library but shouldn't be treated as identifiers.

Parameters:

  • library, version, term, profile (optional)

submit_url_correction

Purpose: Report an incorrect URL detected during scraping (404, redirect loop, login wall). SaddleRAG records the correction and skips the URL in future scrapes.


Configuration tools


list_profiles

Purpose: List all configured MongoDB profiles.


reload_profile

Purpose: Re-read appsettings.json and reconnect to a named profile. Use after manually editing configuration.

Parameters:

  • profile (optional) — reload a specific profile, or all profiles if omitted

Settings tools


set_rerank_strategy

Purpose: Enable or disable cross-encoder reranking at runtime. Takes effect immediately without a server restart.

Parameters:

  • strategy — Off or Onnx
  • profile (optional)

toggle_logging

Purpose: Change the server's log level at runtime.

Parameters:

  • level — Trace, Debug, Information, Warning, Error

Diagnostics tools


get_server_logs

Purpose: Retrieve recent log entries from the server's rolling log file. Useful for diagnosing errors without needing file system access.

Parameters:

  • lines (optional, default 100)
  • level (optional) — filter to a minimum log level

ONNX tools


list_embedding_models

Lists all embedding models registered in appsettings.json with their name, dimensions, and download status.


list_reranker_models

Lists all reranker models registered in appsettings.json.


list_execution_providers

Lists execution providers available on this machine: CPU (always available), DirectML (if GPU detected), CUDA (if NVIDIA GPU with CUDA toolkit detected).


set_active_embedding_model

Switches the active embedding model and writes the selection to runtime-overrides.json. Takes effect on next server restart.

Note: Switching embedding models requires re-embedding all existing libraries with reembed_library since vector spaces are not compatible across models.


set_active_reranker_model

Switches the active reranker model. Pass "none" to disable reranking.


set_execution_provider

Switches between Cpu, DirectMl, and Cuda execution providers. Writes to runtime-overrides.json; takes effect on next restart.


download_onnx_model

Manually trigger a download of a specific ONNX model. By default, model downloads happen automatically at startup; use this to force a re-download after a failed download.

Clone this wiki locally