-
Notifications
You must be signed in to change notification settings - Fork 0
ONNX Models and GPU
SaddleRAG uses ONNX Runtime (ORT) for in-process embedding and reranking. This page covers the models used, how to enable GPU acceleration, model storage, and how to add or switch models.
SaddleRAG ships with two ONNX model configurations:
| Property | Value |
|---|---|
| Model name | nomic-embed-text-v1.5 |
| Architecture | BERT (bi-encoder) |
| Output dimensions | 768 |
| Precision | fp16 |
| File size | ~273 MB |
| Hugging Face repo | nomic-ai/nomic-embed-text-v1.5 |
| Tokenizer | BERT WordPiece (vocab.txt) |
| Task prefixes |
search_document: (indexing), search_query: (querying) |
This model was chosen for its balance of quality, size, and speed. The fp16 export runs efficiently on both CPU and GPU (DirectML/CUDA). The asymmetric task prefix support is critical for good retrieval quality — see Embeddings: ONNX vs Ollama for why this matters.
| Property | Value |
|---|---|
| Model name | mxbai-rerank-base-v1 |
| Architecture | Cross-encoder (sentence pair classifier) |
| Output | Scalar relevance logit |
| Precision | INT8 quantized |
| File size | ~244 MB |
| Hugging Face repo | mixedbread-ai/mxbai-rerank-base-v1 |
| Tokenizer | SentencePiece |
The quantized model trades a small amount of accuracy for ~2–3× faster inference compared to the full-precision model. On a modern CPU, reranking 12 candidates takes ~100–200ms.
Models are stored at:
%ProgramData%\SaddleRAG\models\onnx\
nomic-embed-text-v1.5\
model.onnx
vocab.txt
mxbai-rerank-base-v1\
onnx\
model_quantized.onnx
The location is configurable via Onnx.ModelsDir in appsettings.json.
Model files are downloaded automatically at startup by OnnxModelDownloader. Downloads happen from Hugging Face using the HuggingFaceRepoId and file path configured in appsettings.json. Downloads are atomic — the file is written as {name}.tmp and renamed to the final name only on successful completion, preventing partial file corruption on interrupted downloads.
If a model file already exists, it is not re-downloaded (no hash checking on every startup). To force a re-download, delete the model files and restart the service, or use the download_onnx_model MCP tool.
ONNX Runtime supports multiple hardware backends. SaddleRAG exposes three:
The default. Runs on any x86-64 processor using multi-threaded CPU inference. Thread count is controlled by Onnx.IntraOpNumThreads (0 = auto-detect based on core count).
CPU performance:
- Embedding one chunk: ~5–15ms
- Reranking 12 pairs: ~100–300ms
CPU is fine for light-to-moderate indexing and interactive search. For indexing large documentation sites (hundreds of pages) or frequent batch scrapes, GPU acceleration is significantly faster.
DirectML is Microsoft's hardware-agnostic GPU execution layer for Windows. It works with any modern GPU from any vendor (NVIDIA, AMD, Intel) that supports DirectX 12.
Requirements:
- Windows 10 version 1903 or later
- A DirectX 12-capable GPU
- No separate driver or SDK installation needed — DirectML ships with Windows
DirectML performance (typical modern GPU):
- Embedding one chunk: ~1–3ms
- Reranking 12 pairs: ~20–50ms
For maximum performance on NVIDIA GPUs. Requires:
- NVIDIA GPU with CUDA Compute Capability 6.0+
- CUDA Toolkit 11.x or 12.x installed on the system
CUDA performance is similar to DirectML on NVIDIA hardware; choose based on what's installed on your system.
The installer's CheckGpuCapability custom action queries Win32_VideoController via WMI and pre-selects the appropriate execution provider:
- GPU detected that's not a Microsoft Basic Display Adapter → DirectML pre-selected
- Only Microsoft Basic Display Adapter or Remote Desktop Adapter → CPU pre-selected
- Build older than Windows 10 1903 → installation blocked
If the installer's auto-detection selected CPU but you have a capable GPU, you can switch after installation.
set_execution_provider(provider="DirectMl")
This writes to runtime-overrides.json and takes effect on next server restart.
"Onnx": {
"ExecutionProvider": "DirectMl"
}Valid values: "Cpu", "DirectMl", "Cuda"
list_execution_providers
Returns all providers available on this machine and which one is currently active.
The Onnx.GraphOptimizationLevel setting controls how aggressively ONNX Runtime optimizes the model graph at load time.
| Level | Description |
|---|---|
Disable |
No optimization. Slowest execution. |
Basic |
Constant folding, redundant node removal. Default and recommended. |
Extended |
More complex fusions (layer norm, etc.) |
All |
All available optimizations |
Do not set GraphOptimizationLevel to Extended or All. ONNX Runtime 1.26 contains a SimplifiedLayerNormFusion bug that incorrectly fuses nodes in the fp16 layers of both shipped models, producing incorrect vectors. The Basic level avoids this bug while still providing meaningful speedup over Disable.
You can add any ONNX-exported embedding model that:
-
Accepts
input_idsandattention_maskas inputs -
Outputs hidden states of shape
[batch, sequence, dimensions]that can be mean-pooled -
Uses a BERT WordPiece vocabulary (
vocab.txt) -
Place the
.onnxfile andvocab.txtin%ProgramData%\SaddleRAG\models\onnx\{ModelName}\ -
Add an entry to
Onnx.EmbeddingModelsinappsettings.json:{ "Name": "my-custom-model", "Dimensions": 1024, "HuggingFaceRepoId": "", "ModelFileName": "model.onnx", "VocabFileName": "vocab.txt", "TaskPrefixDocument": "passage: ", "TaskPrefixQuery": "query: " } -
Set
Onnx.ActiveEmbeddingModel = "my-custom-model"(or useset_active_embedding_model) -
Restart the service
-
Re-embed any existing libraries with the new model using
reembed_library
Important: Changing the embedding model requires re-embedding all libraries. Vectors produced by different models are not comparable; searching with a new model against an index built with the old model will return garbage results.
MCP tools for model management:
| Tool | Effect |
|---|---|
list_embedding_models |
Shows all configured models with download status |
list_reranker_models |
Shows all configured reranker models |
set_active_embedding_model(name) |
Switches active embedding model (takes effect on restart) |
set_active_reranker_model(name) |
Switches active reranker model (takes effect on restart) |
set_execution_provider(provider) |
Switches CPU/DirectML/CUDA (takes effect on restart) |
download_onnx_model(name) |
Forces download/re-download of a model |
Settings written by these tools go to runtime-overrides.json, which takes precedence over appsettings.json.