Skip to content

ONNX Models and GPU

Doug Gerard edited this page May 14, 2026 · 1 revision

ONNX Models and GPU Acceleration

SaddleRAG uses ONNX Runtime (ORT) for in-process embedding and reranking. This page covers the models used, how to enable GPU acceleration, model storage, and how to add or switch models.


Default models

SaddleRAG ships with two ONNX model configurations:

Embedding model: nomic-embed-text-v1.5

Property Value
Model name nomic-embed-text-v1.5
Architecture BERT (bi-encoder)
Output dimensions 768
Precision fp16
File size ~273 MB
Hugging Face repo nomic-ai/nomic-embed-text-v1.5
Tokenizer BERT WordPiece (vocab.txt)
Task prefixes search_document: (indexing), search_query: (querying)

This model was chosen for its balance of quality, size, and speed. The fp16 export runs efficiently on both CPU and GPU (DirectML/CUDA). The asymmetric task prefix support is critical for good retrieval quality — see Embeddings: ONNX vs Ollama for why this matters.

Reranker model: mxbai-rerank-base-v1

Property Value
Model name mxbai-rerank-base-v1
Architecture Cross-encoder (sentence pair classifier)
Output Scalar relevance logit
Precision INT8 quantized
File size ~244 MB
Hugging Face repo mixedbread-ai/mxbai-rerank-base-v1
Tokenizer SentencePiece

The quantized model trades a small amount of accuracy for ~2–3× faster inference compared to the full-precision model. On a modern CPU, reranking 12 candidates takes ~100–200ms.


Model files location

Models are stored at:

%ProgramData%\SaddleRAG\models\onnx\
  nomic-embed-text-v1.5\
    model.onnx
    vocab.txt
  mxbai-rerank-base-v1\
    onnx\
      model_quantized.onnx

The location is configurable via Onnx.ModelsDir in appsettings.json.

Automatic download

Model files are downloaded automatically at startup by OnnxModelDownloader. Downloads happen from Hugging Face using the HuggingFaceRepoId and file path configured in appsettings.json. Downloads are atomic — the file is written as {name}.tmp and renamed to the final name only on successful completion, preventing partial file corruption on interrupted downloads.

If a model file already exists, it is not re-downloaded (no hash checking on every startup). To force a re-download, delete the model files and restart the service, or use the download_onnx_model MCP tool.


Execution providers

ONNX Runtime supports multiple hardware backends. SaddleRAG exposes three:

CPU (always available)

The default. Runs on any x86-64 processor using multi-threaded CPU inference. Thread count is controlled by Onnx.IntraOpNumThreads (0 = auto-detect based on core count).

CPU performance:

  • Embedding one chunk: ~5–15ms
  • Reranking 12 pairs: ~100–300ms

CPU is fine for light-to-moderate indexing and interactive search. For indexing large documentation sites (hundreds of pages) or frequent batch scrapes, GPU acceleration is significantly faster.

DirectML (GPU on Windows)

DirectML is Microsoft's hardware-agnostic GPU execution layer for Windows. It works with any modern GPU from any vendor (NVIDIA, AMD, Intel) that supports DirectX 12.

Requirements:

  • Windows 10 version 1903 or later
  • A DirectX 12-capable GPU
  • No separate driver or SDK installation needed — DirectML ships with Windows

DirectML performance (typical modern GPU):

  • Embedding one chunk: ~1–3ms
  • Reranking 12 pairs: ~20–50ms

CUDA (NVIDIA GPU)

For maximum performance on NVIDIA GPUs. Requires:

  • NVIDIA GPU with CUDA Compute Capability 6.0+
  • CUDA Toolkit 11.x or 12.x installed on the system

CUDA performance is similar to DirectML on NVIDIA hardware; choose based on what's installed on your system.


Enabling GPU acceleration

Through the installer

The installer's CheckGpuCapability custom action queries Win32_VideoController via WMI and pre-selects the appropriate execution provider:

  • GPU detected that's not a Microsoft Basic Display Adapter → DirectML pre-selected
  • Only Microsoft Basic Display Adapter or Remote Desktop Adapter → CPU pre-selected
  • Build older than Windows 10 1903 → installation blocked

If the installer's auto-detection selected CPU but you have a capable GPU, you can switch after installation.

At runtime via MCP

set_execution_provider(provider="DirectMl")

This writes to runtime-overrides.json and takes effect on next server restart.

Via appsettings.json

"Onnx": {
  "ExecutionProvider": "DirectMl"
}

Valid values: "Cpu", "DirectMl", "Cuda"

Verify the active provider

list_execution_providers

Returns all providers available on this machine and which one is currently active.


Graph optimization level

The Onnx.GraphOptimizationLevel setting controls how aggressively ONNX Runtime optimizes the model graph at load time.

Level Description
Disable No optimization. Slowest execution.
Basic Constant folding, redundant node removal. Default and recommended.
Extended More complex fusions (layer norm, etc.)
All All available optimizations

Do not set GraphOptimizationLevel to Extended or All. ONNX Runtime 1.26 contains a SimplifiedLayerNormFusion bug that incorrectly fuses nodes in the fp16 layers of both shipped models, producing incorrect vectors. The Basic level avoids this bug while still providing meaningful speedup over Disable.


Adding a custom embedding model

You can add any ONNX-exported embedding model that:

  1. Accepts input_ids and attention_mask as inputs

  2. Outputs hidden states of shape [batch, sequence, dimensions] that can be mean-pooled

  3. Uses a BERT WordPiece vocabulary (vocab.txt)

  4. Place the .onnx file and vocab.txt in %ProgramData%\SaddleRAG\models\onnx\{ModelName}\

  5. Add an entry to Onnx.EmbeddingModels in appsettings.json:

    {
      "Name": "my-custom-model",
      "Dimensions": 1024,
      "HuggingFaceRepoId": "",
      "ModelFileName": "model.onnx",
      "VocabFileName": "vocab.txt",
      "TaskPrefixDocument": "passage: ",
      "TaskPrefixQuery": "query: "
    }
  6. Set Onnx.ActiveEmbeddingModel = "my-custom-model" (or use set_active_embedding_model)

  7. Restart the service

  8. Re-embed any existing libraries with the new model using reembed_library

Important: Changing the embedding model requires re-embedding all libraries. Vectors produced by different models are not comparable; searching with a new model against an index built with the old model will return garbage results.


Switching models at runtime

MCP tools for model management:

Tool Effect
list_embedding_models Shows all configured models with download status
list_reranker_models Shows all configured reranker models
set_active_embedding_model(name) Switches active embedding model (takes effect on restart)
set_active_reranker_model(name) Switches active reranker model (takes effect on restart)
set_execution_provider(provider) Switches CPU/DirectML/CUDA (takes effect on restart)
download_onnx_model(name) Forces download/re-download of a model

Settings written by these tools go to runtime-overrides.json, which takes precedence over appsettings.json.

Clone this wiki locally