Repository navigation
v20260118: CUDA!
·
4 commits
to main
since this release
nomic-embed-v1.5
Multimodal embedding server supporting both text and image embeddings.
Installation
Important: You need to download both a binary and a models package, then combine them.
Step 1: Download Binary
Download the binary for your platform:
# Linux x86_64
tar -xzf nomic-embed-v1.5-20260118-linux-x64.tar.gz
# Linux ARM64 (Raspberry Pi, AWS Graviton, Apple Silicon Linux VMs)
tar -xzf nomic-embed-v1.5-20260118-linux-arm64.tar.gz
# macOS (Apple Silicon M1/M2/M3)
tar -xzf nomic-embed-v1.5-20260118-macos-arm64.tar.gzStep 2: Download Models
Download one of the model packages (choose based on your needs):
# Quantized models (smaller, faster, ~99% quality) - Recommended for most use cases
tar -xzf nomic-embed-v1.5-20260118-models-quantized.tar.gz
# OR Full precision models (larger, best quality, fp32 precision)
tar -xzf nomic-embed-v1.5-20260118-models-full.tar.gzStep 3: Combine Binary and Models
Extract both archives into the same directory:
# Example: Extract both to the same location
mkdir -p nomic-embed-v1.5-20260118
cd nomic-embed-v1.5-20260118
# Extract binary (creates nomic-embed-v1.5-20260118-<platform>/nomic-serve)
tar -xzf ../nomic-embed-v1.5-20260118-linux-x64.tar.gz
# Extract models (creates nomic-embed-v1.5-20260118-models-<variant>/models/)
tar -xzf ../nomic-embed-v1.5-20260118-models-quantized.tar.gz
# Move binary and models to the same directory
mv nomic-embed-v1.5-20260118-linux-x64/nomic-serve .
mv nomic-embed-v1.5-20260118-models-quantized/models .Alternative (simpler): Extract both archives and copy files manually:
# Extract binary
tar -xzf nomic-embed-v1.5-20260118-linux-x64.tar.gz
# Extract models
tar -xzf nomic-embed-v1.5-20260118-models-quantized.tar.gz
# Copy models into binary directory
cp -r nomic-embed-v1.5-20260118-models-quantized/models \
nomic-embed-v1.5-20260118-linux-x64/
# Your final structure should be:
# nomic-embed-v1.5-20260118-linux-x64/
# ├── nomic-serve
# └── models/
# ├── txt/
# │ ├── model_quantized.onnx (or model.onnx)
# │ └── tokenizer.json
# └── img/
# └── model_quantized.onnx (or model.onnx)Running
cd nomic-embed-v1.5-20260118-<platform>
./nomic-serve
# Server starts on http://localhost:8080The server will automatically detect which model files are present.
If both model.onnx and model_quantized.onnx exist, it will prefer the full precision model.
You can override this by setting TXT_MODEL and IMG_MODEL environment variables explicitly.
API Endpoints
POST /txt/embed- Single text embeddingPOST /txt/batch- Batch text embeddingsPOST /txt/query- Query with enforced search_query prefixPOST /img/embed- Single image embeddingPOST /img/batch- Batch image embeddingsGET /health- Health checkGET /info- Server information (model paths, configuration)GET /docs- Swagger UI
Environment Variables
PORT- Server port (default: 8080)TXT_MODEL- Path to text ONNX model (optional, auto-detectsmodel.onnxormodel_quantized.onnx)TOKENIZER- Path to tokenizer file (default: models/txt/tokenizer.json)IMG_MODEL- Path to vision ONNX model (optional, auto-detectsmodel.onnxormodel_quantized.onnx)AVERAGING- Averaging method for image color statistics:arithmeticorgeometric(default:geometric)
Model Variants:
- Quantized (
models-quantized.tar.gz): ~224MB total (text: 131MB, vision: 92MB, tokenizer: 696KB), faster inference, ~99% quality. Recommended for most use cases. - Full (
models-full.tar.gz): ~880MB total (text: 522MB, vision: 357MB, tokenizer: 696KB), best quality (fp32 precision). Use for maximum accuracy.
What's Changed
- cuda compatibility + overall improvements by @mathematicalmichael in #2
New Contributors
- @mathematicalmichael made their first contribution in #2
Full Changelog: v20260106...v20260118