Skip to content

Why SaddleRAG

Doug Gerard edited this page May 14, 2026 · 1 revision

Why SaddleRAG

The problem

AI coding assistants are excellent at reasoning about code. But their knowledge is frozen at a training cutoff. When you ask about a library that was updated after that cutoff — or a library niche enough to be underrepresented in training data — the assistant either admits it doesn't know, or worse, confidently generates plausible-sounding but incorrect code.

This shows up in everyday work as:

  • Wrong API signatures. The assistant generates a method call that existed in version 2.x but was changed in version 3.x.
  • Missing features. A new configuration option or helper that the library added last year is unknown to the assistant.
  • Invented APIs. The assistant "knows" the general shape of the library and fills in specifics that don't exist — a phenomenon sometimes called hallucination.
  • Internal documentation blind spots. Your organization's internal libraries, toolkits, and APIs don't appear in any public training dataset.

Why not just paste the docs into the context?

The context window approach has hard limits:

  • Documentation for even a moderately sized library runs into hundreds of pages. API reference docs alone often exceed 100,000 tokens, far beyond what fits in a single context window.
  • Pasting docs manually is tedious and doesn't persist across conversations.
  • Most documentation isn't available as plain text — it's rendered HTML across dozens or hundreds of URLs.

Why not use web search?

Web search helps with popular, public libraries that have well-indexed documentation. It doesn't help when:

  • Your library is internal (not on the internet)
  • You want to search a specific version of the docs, not whatever Google returns
  • You need sub-page retrieval (find the one class method out of hundreds)
  • Search quality varies and irrelevant results can mislead the assistant

The RAG approach

Retrieval-Augmented Generation (RAG) addresses this by giving the assistant a retrieval mechanism it can call mid-conversation. Instead of baking all knowledge into the model, you store knowledge externally and look it up on demand.

The pattern looks like this:

User asks: "How do I configure retry policies in Polly 8?"

Without RAG:  assistant guesses from training data (possibly wrong or stale)

With RAG:
  1. assistant embeds the question as a vector
  2. queries the index for the closest chunks from the Polly 8 docs
  3. receives the actual relevant sections
  4. answers the user using that retrieved context

RAG systems built for open-domain question answering (Wikipedia, general web crawls) are common. RAG for specific library documentation — with version control, structured chunking, and type-aware retrieval — is what SaddleRAG provides.


Why documentation is different from general RAG

Generic RAG systems treat all text the same way: split at fixed token counts, embed, store, retrieve. That works fine when all your text is similar in structure (e.g., news articles). Documentation is not like that.

A documentation site contains fundamentally different kinds of pages:

Page type What it contains What breaks if you split it wrong
API reference Class definitions, method signatures, parameter tables Splitting in the middle of a method entry makes the chunk useless
Code samples Working code examples Splitting a code block at a line boundary breaks the example
HowTo guides Procedural steps with explanations Splitting mid-procedure loses causality
Changelogs Version-by-version release notes Mixing version entries in one chunk makes version-specific answers impossible
Overview Architecture explanations, conceptual diagrams Can tolerate section splits

SaddleRAG classifies every page before chunking it, then applies a strategy appropriate to the category. A code sample page is kept whole. A changelog is split at version boundaries. An API reference is split at class boundaries. This category-aware chunking is one of the core reasons SaddleRAG produces better retrieval results than a generic RAG pipeline.


The team dimension

The problem multiplies for teams. If every developer runs their own AI assistant independently:

  • Each developer must manually keep their own assistant updated when libraries change
  • No developer benefits from documentation that a colleague already indexed
  • CI/CD pipelines can't pre-index documentation for newly adopted libraries before the team needs them

SaddleRAG supports a shared deployment model where a team runs one SaddleRAG instance against a shared MongoDB database. Any developer's AI assistant can query the shared index. A CI/CD pipeline can trigger indexing for new library versions automatically. See Profiles and Team Usage for how this works.


Why not an existing solution?

The tools that existed when SaddleRAG was designed either:

  • Required cloud infrastructure (vector databases, cloud embedding APIs) adding cost, complexity, and data egress concerns
  • Were general-purpose and not designed for the structured heterogeneity of documentation sites
  • Had no version awareness — they indexed a library once and couldn't track when docs changed
  • Had no category-aware chunking — API reference entries got the same treatment as blog posts
  • Required MongoDB Atlas for vector search (SaddleRAG uses in-process in-memory vector search, no Atlas required)

SaddleRAG is specifically designed to solve the "keep my AI assistant current on the exact library versions my project uses" problem, locally, without cloud dependencies.

Clone this wiki locally