-
Notifications
You must be signed in to change notification settings - Fork 0
Page Classification
Every page crawled by SaddleRAG is classified into one of seven categories before it is chunked. This classification step is one of SaddleRAG's most important design choices — it is what separates documentation-aware RAG from generic text chunking.
The fundamental problem is that documentation pages are not uniform. A single documentation site contains pages that are structurally very different from each other:
- A conceptual overview page flows as paragraphs of prose
- An API reference page is a structured list of classes, methods, and parameters
- A code sample page is primarily a code block with brief surrounding explanation
- A changelog page is a timestamped list of versioned changes
If you split all of these with the same strategy (e.g., "split every 512 tokens"), you get:
- Code examples split mid-function, producing unusable fragments
- API reference entries split so that a method signature is in one chunk and its description is in another
- Changelog entries mixed across version boundaries, making version-specific queries impossible
Classification lets the chunker apply the right strategy for each page type. See Ingestion Pipeline — Stage 3: Chunk for the full chunking strategies.
Classification also drives the category filter in search_docs. You can restrict a search to ApiReference content only (useful when you want exact method signatures) or Sample content only (useful when you want working code examples).
What it is: Conceptual, architectural, or introductory content. These pages explain what something is and why it exists, not how to use it step by step.
Examples: "Introduction to Polly," "Architecture overview," "Core concepts," "Design philosophy"
Chunking: Split at heading boundaries. Each heading + following content becomes a chunk.
What it is: Procedural, tutorial, or guide content. These pages explain how to accomplish a specific task, usually with numbered steps or sequential instructions.
Examples: "How to configure retry policies," "Getting started guide," "Tutorial: building a resilience pipeline," "Migrating from v2 to v3"
Chunking: Split at heading boundaries.
What it is: Pages whose primary content is one or more self-contained code examples. The explanation is secondary to the code.
Examples: "Code samples," "Examples," standalone snippet pages, "Recipes"
Chunking: Kept as a single chunk, regardless of length. Code examples must not be split.
What it is: Actual source files, typically from a GitHub repository that was indexed rather than a documentation site. The entire file is the content.
Examples: Any .cs, .ts, .py, .go source file indexed via the GitHub repo scraper
Chunking: Kept as a single chunk. Source files are treated atomically; splitting in the middle of a class or function produces unusable fragments.
What it is: API documentation — class definitions, method signatures, property descriptions, parameter tables. These pages are the authoritative reference for what an API accepts and returns.
Examples: "ResiliencePipeline class," "RetryStrategyOptions properties," "AddRetry method," auto-generated API docs from DocFX, Swagger/OpenAPI pages
Chunking: Split at heading boundaries. For well-structured API docs, each heading typically corresponds to one class or one method group, so each chunk is a coherent API reference unit.
What it is: Release notes, version history, or migration guides. These pages list changes by version.
Examples: "Changelog," "Release notes," "What's new," "Migration guide," "Breaking changes"
Chunking: Split at version boundary markers. The chunker recognizes patterns like ## 8.2.0, # v3.0, ### 2024-01-15 and makes each version entry its own chunk.
What it is: Pages where classification failed, timed out, or returned a confidence score below the threshold.
Chunking: Falls back to heading-boundary splitting (same as Overview). This is a safe default — it preserves heading structure without making assumptions about content type.
Implementation: SaddleRAG.Ingestion/Classification/LlmClassifier.cs
For each page, the classifier:
- Extracts the page title, URL, and the first 500 characters of text content
- Sends a prompt to Ollama asking for a JSON response with
categoryandconfidencefields - Streams the response and parses the JSON
- If JSON parsing fails, scans the raw response text for a
DocCategoryenum name as a fallback - Updates the
PageRecordwith the category and confidence
The prompt is a zero-shot instruction that:
- Lists all seven valid category names with a one-line description of each
- Provides the library name, page URL, page title, and content preview
- Requests JSON output in a specific schema
- Instructs the model to default to
Unclassifiedwhen uncertain
The PromptVersion constant in the source code tracks when the prompt changes, so that a re-classification run knows which pages were classified with an outdated prompt.
The active classification model is resolved from Ollama.ClassificationModels — an ordered list of model entries. The first entry is used unless Ollama.ActiveClassificationModel names a different one.
The default model is phi4-mini:3.8b. This is a 3.8-billion-parameter model from Microsoft. At this size it:
- Runs fast on CPU (a few hundred milliseconds per classification)
- Has strong instruction-following ability (reliably produces valid JSON)
- Understands documentation structure well enough to classify accurately
- Requires ~2.4 GB of VRAM/RAM when loaded by Ollama
Classification failures are non-fatal. If Ollama is unreachable, the model doesn't respond, or the response can't be parsed, the page is assigned Unclassified and pipeline processing continues. A failed classification is logged and counted in the scrape audit log.
If Ollama is down for an entire scrape run, all pages will be Unclassified. The scrape will succeed, but search quality will be reduced (because chunking defaults to heading-split for everything). You can re-run classification later with rechunk_library after Ollama is restored.
Classification is more accurate when the classifier knows about the library's naming conventions, language, and URL structure. The Library Reconnaissance step, which runs before scraping, produces a LibraryProfile that includes:
- The programming language(s) the library targets
- Naming convention hints (e.g., PascalCase class names, camelCase methods)
- Known symbols and terms to watch for
This profile is available to the classifier during the scrape. The library name provided in the prompt helps the model apply appropriate priors — a page called "Stream" in a .NET library is probably API reference; in a streaming media library it might be an overview.
To switch to a different Ollama model for classification:
- Pull the model:
ollama pull mistral(or whichever model you want) - Add it to
appsettings.json:"Ollama": { "ClassificationModels": [ { "Name": "mistral" }, { "Name": "phi4-mini" } ], "ActiveClassificationModel": "mistral" }
- Restart SaddleRAGMcp
- Re-run classification: use
rechunk_librarywithreclassify: trueor re-scrape from scratch
Any model that follows instructions and can produce valid JSON output works. Larger models may classify edge cases more accurately; smaller models are faster. For documentation classification — a relatively structured task with seven well-defined options — models in the 3B–8B range perform well.
There are two distinct Ollama model roles in SaddleRAG:
| Role | Setting | Default | When used |
|---|---|---|---|
| Recon model | Ollama.ReconModels |
phi4-mini:3.8b |
During recon_library: generates the LibraryProfile (language, casing conventions, likely symbols) when a frontier LLM isn't available |
| Classification model | Ollama.ClassificationModels |
phi4-mini:3.8b |
During ingestion Stage 2: classifies each page |
These can be configured independently. The recon model does a more complex task (open-ended library analysis) and may benefit from a larger model. The classification model does a simpler, more structured task and a small model works well.