Repository navigation
Ingestion Pipeline
The ingestion pipeline takes a documentation URL and produces a fully indexed, searchable knowledge base. It runs in the background after a scrape_docs MCP tool call or a saddlerag ingest CLI command.
The pipeline is structured as five stages connected by bounded producer/consumer channels. Each stage runs concurrently in its own loop; the channels provide natural back-pressure so a slow stage (e.g., Ollama classification) doesn't cause the faster stages ahead of it to buffer unboundedly in RAM.
┌─────────────┐
URL ──► │ Stage 1 │ PageCrawler
│ Crawl │ Playwright BFS
└──────┬──────┘ Channel capacity: 50 pages
│
┌──────▼──────┐
│ Stage 2 │ LlmClassifier
│ Classify │ Ollama phi4-mini
└──────┬──────┘
│
┌──────▼──────┐
│ Stage 3 │ CategoryAwareChunker
│ Chunk │ + SymbolExtractor
└──────┬──────┘ Channel capacity: 20 chunk lists
│
┌──────▼──────┐
│ Stage 4 │ OnnxEmbeddingProvider
│ Embed │ (or OllamaEmbeddingProvider)
└──────┬──────┘ Batch size: 32 chunks
│
┌──────▼──────┐
│ Stage 5 │ InMemoryVectorSearch
│ Index │ + MongoDB upsert
└─────────────┘
│
Post-pipeline
├── Build BM25 index
├── Update LibraryRecord metadata
└── Run SuspectDetector
Implementation: SaddleRAG.Ingestion/Crawling/PageCrawler.cs
The crawler uses a Playwright headless Chromium browser to fetch pages. This is deliberate: many modern documentation sites render their content via JavaScript. A plain HTTP fetcher would get empty pages. Playwright executes the page's JavaScript and waits for the DOM to settle before extracting content.
The crawler uses breadth-first search with three priority channels:
-
InScopeEntries — URLs that match the root documentation path prefix (e.g.,
https://docs.example.com/api/). These are drained first. - OffPathEntries — URLs that were linked from in-scope pages but don't share the root path prefix (e.g., a link to the company blog). These are crawled only after in-scope pages are exhausted, and only if the crawl budget allows.
- RetryEntries — Pages that failed transiently (network error, timeout). Retried with backoff.
This priority structure ensures that if you point SaddleRAG at https://docs.example.com/api/, it thoroughly covers the API docs before following off-path links.
Two pattern lists (from LibraryProfile) control what gets crawled:
-
AllowedUrlPatterns— regex patterns; only URLs matching at least one are fetched -
ExcludedUrlPatterns— regex patterns; URLs matching any of these are skipped
If AllowedUrlPatterns is empty, all URLs reachable from the seed URL are eligible (subject to the same-origin and root-path constraints).
The CrawlBudget setting caps the total number of pages fetched. The default is generous (several thousand pages), but for very large documentation sites you may want to narrow scope using URL patterns rather than relying solely on the budget.
If the seed URL is a GitHub repository (matches the GitHub URL pattern), the crawler delegates to GitHubRepoScraper instead of Playwright. This clones the repository and walks the file tree, treating each .md, .txt, .rst, and source file as a "page." This is useful for indexing documentation that lives in the repo itself (README, docs/ folder, inline source comments).
If a scrape was interrupted, start_ingest will detect the existing job and offer to resume. The crawler skips URLs already present in the pages collection (matched by LibraryId, Version, and URL). New or changed URLs are crawled fresh.
Implementation: SaddleRAG.Ingestion/Classification/LlmClassifier.cs
Classification is the step that makes SaddleRAG's chunking intelligent. Every page is sent to Ollama with a structured prompt requesting a category and confidence score.
The classifier sends:
- The library name (for context)
- The page URL (the path often reveals the content type, e.g.,
/api/vs/guides/) - The page title
- Up to 500 characters of extracted text content (the preview)
And asks for a JSON response in the form:
{ "category": "ApiReference", "confidence": 0.92 }| Category | What it covers |
|---|---|
Overview |
Architecture explanations, conceptual introductions, "getting started" narrative pages |
HowTo |
Step-by-step guides, tutorials, walkthroughs with procedural instructions |
Sample |
Standalone code samples and demos whose primary content is working code |
Code |
Source files (from GitHub repos) — functions, classes, modules |
ApiReference |
Class/method/property/parameter documentation |
ChangeLog |
Release notes, version history, migration guides |
Unclassified |
Pages where classification failed or confidence was too low |
The active classification model is the first entry in Ollama.ClassificationModels unless Ollama.ActiveClassificationModel names a specific one. The default is phi4-mini:3.8b — a 3.8-billion-parameter model that classifies accurately with low latency.
Classification failures are non-fatal. If Ollama is unreachable, returns malformed JSON, or the model produces output that doesn't parse, the page proceeds as Unclassified. An Unclassified page is chunked at heading boundaries — the same as Overview — which is a reasonable fallback.
The alternatives are:
- Heuristics (URL patterns, title keywords) — work for well-structured docs but fail when a page titled "Overview" is actually an API reference, or when URL paths don't follow conventions.
- Frontier LLM API — accurate, but adds latency, cost, and requires sending your documentation content to an external service.
- Local LLM via Ollama — accurate enough for this binary-ish classification task, zero cost, no data egress, runs on the same machine.
At 3.8B parameters, phi4-mini classifies a 500-character preview reliably for the seven categories. The latency per page is a few hundred milliseconds, which is dominated by the Playwright crawl time anyway.
Implementation: SaddleRAG.Ingestion/Chunking/CategoryAwareChunker.cs
The chunker splits each page's content into retrieval-sized pieces. The strategy depends on category.
Sample and Code — kept as a single chunk. Code examples must not be split; splitting a code block in the middle produces a useless fragment. These pages are treated as atomic units regardless of length.
HowTo, ApiReference, Overview, Unclassified — split at Markdown heading boundaries (#, ##, ###). Each heading and the content that follows it becomes a chunk. This preserves the logical structure of the page: a heading introduces a topic, and the content below it elaborates on that topic — they belong together.
ChangeLog — split at version boundary markers. The chunker recognizes patterns like ## 8.2.0, # v3.0, ### 2024-01-15 and treats each version entry as one chunk. This means a query like "what changed in 8.2.0?" retrieves exactly the 8.2.0 entry rather than a fragment of it alongside entries from surrounding versions.
After the category-aware primary split, any chunk that exceeds MaxChunkChars is further split at sentence boundaries. The chunker walks backward from the size limit to find a sentence end (., ?, ! followed by whitespace). This prevents oversized chunks from degrading embedding quality (most embedding models have a 512-token input limit).
After chunking, SymbolExtractor scans each chunk for identifier-shaped tokens: sequences of alphanumeric characters and underscores that match the casing conventions documented in the library's LibraryProfile. These become the chunk's Symbols list and its primary QualifiedName if one can be determined.
The Symbols list is used at search time for the "symbol backstop" — a secondary retrieval path that finds chunks by identifier name when the vector search misses them.
Implementation: SaddleRAG.Ingestion/Embedding/OnnxEmbeddingProvider.cs (default)
Embedding converts each chunk's text into a dense vector representation. Two chunks about similar topics will have similar vectors; dissimilar topics will be far apart in vector space.
The raw chunk content is not embedded directly. The embedding input is:
[ApiReference] [polly] [ResiliencePipeline class]
AddRetry configures retry resilience strategy...
The [Category] [LibraryId] [PageTitle] prefix is prepended to every chunk. This helps the model place the chunk in the correct semantic neighborhood — API reference content about "AddRetry" in "Polly" is distinct from, say, a HowTo guide that mentions retries in a different library.
Additionally, the ONNX embedding provider applies a document task prefix ("search_document: ") before this enriched text. This is the asymmetric bi-encoder feature: the same nomic-embed-text-v1.5 model is used for both indexing and querying, but with different task prefixes. Documents are prefixed with search_document: at index time; queries are prefixed with search_query: at query time. The model was trained on this asymmetric scheme, and using it correctly is essential for good retrieval quality.
Chunks arrive at Stage 4 in batches of 32 (the EmbedBatchSize setting). The ONNX provider processes them serially within the batch, but the pipeline accumulates chunks before persisting — so MongoDB upserts happen in groups rather than one-by-one.
Each chunk is upserted to MongoDB immediately after embedding. This supports resume: if a scrape job is interrupted, a restarted scrape will find the already-embedded chunks in MongoDB and skip re-embedding them.
Implementation: SaddleRAG.Ingestion/Embedding/InMemoryBruteForceVectorSearch.cs
The final stage adds the embedded chunks to the in-memory vector index. The index is a dictionary keyed by {profile}/{libraryId}/{version}, holding a List<DocChunk>. Cosine similarity search runs over this list at query time.
A ScrapeAuditLogEntry is written to MongoDB for every page processed — recording the URL, HTTP status, category, chunk count, and skip reason (if applicable). This audit log drives the get_library_health diagnostics tool.
After all five stages drain:
Bm25IndexBuilder walks every chunk for the (library, version) and builds a sharded inverted index. Each shard covers a partition of the chunk id space and stores term → {chunkId, tf-idf weight} mappings. The shard data is stored in MongoDB's bm25Shards collection (with a GridFS fallback for oversized shards). A LibraryIndex document is upserted with the corpus statistics (document count, average document length, term count) needed for BM25 scoring at query time.
LibraryRecord and LibraryVersionRecord are updated with the completed scrape's page count, chunk count, and the name of the embedding provider and model used. The embedding provider name ("onnx" or "ollama") and model name are persisted so that future maintenance operations (re-embed, re-chunk) and the startup bootstrapper know which model family to use for this version.
SuspectDetector runs heuristics on the completed scrape results to identify suspicious outcomes:
- Very few pages indexed from a URL that should have many
- Very high fraction of
Unclassifiedpages - All pages mapping to a single category (suggests the URL may be a 404 page or a login wall)
- Very low total chunk count relative to page count
Libraries flagged as suspect appear in get_dashboard_index output and are called out by start_ingest when a re-scrape is requested.
A scrape job progresses through these states:
Queued → Running → Completed
→ Failed
→ Cancelled
Jobs are stored in the scrapeJobs collection with a 30-day TTL. The scrape_audit_log collection records per-page outcomes for the most recent job per library version, also with a 30-day TTL.
Use get_scrape_status (MCP) or saddlerag status (CLI) to query job state. Use cancel_scrape to abort a running job.