Skip to content

Library Reconnaissance

Doug Gerard edited this page May 14, 2026 · 1 revision

Library Reconnaissance

Before SaddleRAG indexes a new library, it runs a reconnaissance step to learn about the library's structure, conventions, and terminology. The output of this step — the LibraryProfile — shapes every downstream stage of the pipeline.


What is the LibraryProfile?

The LibraryProfile is a structured document that captures library-specific knowledge that cannot be inferred from documentation text alone. It includes:

  • Programming language(s) the library targets
  • Identifier casing conventions (PascalCase classes? camelCase methods? UPPER_SNAKE_CASE constants?)
  • Known symbols — class names, method names, and other identifiers the library exposes
  • URL patterns — which URL paths contain real documentation vs. navigation pages, changelogs, etc.
  • Stoplist — common words in this library's docs that look like identifiers but aren't (e.g., the word "Configuration" in a general-purpose library)
  • LikelySymbols — identifiers the symbol extractor should specifically look for

This information is used by:

  • Stage 2 (Classification): The library name and language hint help the LLM classify pages more accurately
  • Stage 3 (Chunking): URL patterns drive the AllowedUrlPatterns and ExcludedUrlPatterns for the crawl; casing conventions guide the symbol extractor
  • Stage 3 (Symbol extraction): LikelySymbols and Stoplist tune what gets extracted as an identifier

How recon works

Preferred path: frontier LLM

The start_ingest state machine, when it determines that a LibraryProfile doesn't yet exist for the requested library/version, returns RECON_NEEDED with detailed instructions for the calling AI assistant. It provides the calling LLM with:

  1. The full LibraryProfile JSON schema and field descriptions
  2. Instructions to use its training knowledge about the library to fill in the profile
  3. The URL to be scraped and any other context available

The calling LLM (e.g., Claude) fills in the profile using its training data. For well-known libraries (Polly, Serilog, MassTransit, React, etc.), the LLM already knows the API surface, naming conventions, and URL structure. The filled-in profile is submitted back via recon_library or submit_library_profile.

Why this is the preferred path: Frontier LLMs (Claude, GPT-4) have extensive knowledge about popular libraries. They can fill in a high-quality LibraryProfile in seconds, without any scraping. This knowledge complements what SaddleRAG will discover during the actual scrape.

Fallback path: local Ollama recon

If the calling LLM is not available or is running in an offline/restricted environment, SaddleRAG can generate the LibraryProfile using a local Ollama model. The CLI's recon command triggers this fallback:

saddlerag recon --url https://docs.example.com/ --library example-lib --version 1.0.0

The Ollama recon model (Ollama.ReconModels, default phi4-mini) is given the library URL and name and asked to produce a JSON LibraryProfile. The output is validated against Ollama.ReconMinConfidence (default 0.6) — if the model's confidence is too low, the profile is not persisted.

Ollama recon is less accurate than frontier LLM recon for well-known libraries (because smaller models have less breadth of knowledge), but it works well for libraries where the URL structure and language are clear.

Manual profile submission

The submit_library_profile MCP tool accepts a complete LibraryProfile JSON and persists it directly. This is useful for:

  • Libraries not well-known to any LLM (internal libraries)
  • Overriding a generated profile with manually curated data
  • Batch setup scripts that pre-populate profiles before scraping

The LibraryProfile schema

{
  "LibraryId": "polly",
  "Version": "8.2.0",
  "Language": "csharp",
  "NamingConventions": {
    "ClassCasing": "PascalCase",
    "MethodCasing": "PascalCase",
    "PropertyCasing": "PascalCase",
    "ConstantCasing": "PascalCase"
  },
  "LikelySymbols": [
    "ResiliencePipeline",
    "RetryStrategyOptions",
    "CircuitBreakerStrategyOptions",
    "AddRetry",
    "AddCircuitBreaker",
    "AddTimeout",
    "AddHedging"
  ],
  "Stoplist": [
    "configuration",
    "strategy",
    "options"
  ],
  "AllowedUrlPatterns": [
    "pollydocs\\.org/docs/",
    "pollydocs\\.org/api/"
  ],
  "ExcludedUrlPatterns": [
    "pollydocs\\.org/blog/",
    "pollydocs\\.org/community/"
  ],
  "DocRootUrl": "https://www.pollydocs.org/",
  "Confidence": 0.95,
  "Notes": "Polly 8 uses ResiliencePipeline instead of v7's Policy classes"
}

Field descriptions

Field Description
LibraryId The library identifier (matches the library parameter in tool calls)
Version The version this profile applies to
Language Primary programming language: csharp, typescript, javascript, python, go, java, rust, etc.
NamingConventions Casing conventions for different identifier types
LikelySymbols Known exported identifiers. The symbol extractor prioritizes finding these in chunk content. These are also returned by list_symbols.
Stoplist Common words in this library's documentation that look like identifiers but aren't. These are excluded from BM25 indexing and symbol extraction.
AllowedUrlPatterns Regex patterns. Only URLs matching at least one are fetched. If empty, all reachable URLs within the same root are fetched.
ExcludedUrlPatterns Regex patterns. URLs matching any of these are skipped regardless of AllowedUrlPatterns.
DocRootUrl The canonical root URL for this library's documentation
Confidence 0.0–1.0 confidence score in the profile's accuracy (used by Ollama recon to filter low-quality results)
Notes Free-text notes about the library or version (included in the classification prompt for context)

How the profile improves results

Better URL filtering

Without a profile, SaddleRAG crawls everything reachable from the seed URL. For most documentation sites this works, but some sites have:

  • A blog section that's irrelevant to API users
  • Community forums linked from the sidebar
  • Multiple product versions at different URL paths

AllowedUrlPatterns lets you say "only index URLs under /docs/ and /api/." The crawler skips everything else, producing a cleaner, faster, more relevant index.

Better symbol extraction

The LikelySymbols list tells the symbol extractor what to look for. When a chunk mentions ResiliencePipeline.AddRetry, the extractor knows this is an intentional API reference (it's in LikelySymbols) rather than incidental text.

The Stoplist prevents false positives. If "Pipeline" appears as a generic word in many chunks (not the Polly-specific ResiliencePipeline), adding "pipeline" to the stoplist prevents it from being extracted as a symbol everywhere.

Better classification context

The classification prompt includes the library name and Notes from the profile. This gives the LLM classifier additional context. A note like "Polly 8 uses ResiliencePipeline instead of v7's Policy classes" helps the classifier correctly categorize pages that discuss this migration.


Updating a profile

To update an existing profile after discovering issues (wrong URL patterns, missing symbols):

  1. Use submit_library_profile to push a corrected profile
  2. Run reextract_library to re-run symbol extraction with the updated LikelySymbols and Stoplist
  3. If the URL patterns changed significantly, consider running rescrape_library to crawl with the corrected patterns

Alternatively, use rechunk_library with reclassify: true to re-run both classification and chunking with the updated profile in one operation.

Clone this wiki locally