Skip to content

v20260118: CUDA!

Choose a tag to compare

@mathematicalmichael mathematicalmichael released this 18 Jan 23:27
· 4 commits to main since this release

nomic-embed-v1.5

Multimodal embedding server supporting both text and image embeddings.

Installation

Important: You need to download both a binary and a models package, then combine them.

Step 1: Download Binary

Download the binary for your platform:

# Linux x86_64
tar -xzf nomic-embed-v1.5-20260118-linux-x64.tar.gz

# Linux ARM64 (Raspberry Pi, AWS Graviton, Apple Silicon Linux VMs)
tar -xzf nomic-embed-v1.5-20260118-linux-arm64.tar.gz

# macOS (Apple Silicon M1/M2/M3)
tar -xzf nomic-embed-v1.5-20260118-macos-arm64.tar.gz

Step 2: Download Models

Download one of the model packages (choose based on your needs):

# Quantized models (smaller, faster, ~99% quality) - Recommended for most use cases
tar -xzf nomic-embed-v1.5-20260118-models-quantized.tar.gz

# OR Full precision models (larger, best quality, fp32 precision)
tar -xzf nomic-embed-v1.5-20260118-models-full.tar.gz

Step 3: Combine Binary and Models

Extract both archives into the same directory:

# Example: Extract both to the same location
mkdir -p nomic-embed-v1.5-20260118
cd nomic-embed-v1.5-20260118

# Extract binary (creates nomic-embed-v1.5-20260118-<platform>/nomic-serve)
tar -xzf ../nomic-embed-v1.5-20260118-linux-x64.tar.gz

# Extract models (creates nomic-embed-v1.5-20260118-models-<variant>/models/)
tar -xzf ../nomic-embed-v1.5-20260118-models-quantized.tar.gz

# Move binary and models to the same directory
mv nomic-embed-v1.5-20260118-linux-x64/nomic-serve .
mv nomic-embed-v1.5-20260118-models-quantized/models .

Alternative (simpler): Extract both archives and copy files manually:

# Extract binary
tar -xzf nomic-embed-v1.5-20260118-linux-x64.tar.gz

# Extract models
tar -xzf nomic-embed-v1.5-20260118-models-quantized.tar.gz

# Copy models into binary directory
cp -r nomic-embed-v1.5-20260118-models-quantized/models \
      nomic-embed-v1.5-20260118-linux-x64/

# Your final structure should be:
# nomic-embed-v1.5-20260118-linux-x64/
#   ├── nomic-serve
#   └── models/
#       ├── txt/
#       │   ├── model_quantized.onnx (or model.onnx)
#       │   └── tokenizer.json
#       └── img/
#           └── model_quantized.onnx (or model.onnx)

Running

cd nomic-embed-v1.5-20260118-<platform>
./nomic-serve
# Server starts on http://localhost:8080

The server will automatically detect which model files are present.
If both model.onnx and model_quantized.onnx exist, it will prefer the full precision model.
You can override this by setting TXT_MODEL and IMG_MODEL environment variables explicitly.

API Endpoints

  • POST /txt/embed - Single text embedding
  • POST /txt/batch - Batch text embeddings
  • POST /txt/query - Query with enforced search_query prefix
  • POST /img/embed - Single image embedding
  • POST /img/batch - Batch image embeddings
  • GET /health - Health check
  • GET /info - Server information (model paths, configuration)
  • GET /docs - Swagger UI

Environment Variables

  • PORT - Server port (default: 8080)
  • TXT_MODEL - Path to text ONNX model (optional, auto-detects model.onnx or model_quantized.onnx)
  • TOKENIZER - Path to tokenizer file (default: models/txt/tokenizer.json)
  • IMG_MODEL - Path to vision ONNX model (optional, auto-detects model.onnx or model_quantized.onnx)
  • AVERAGING - Averaging method for image color statistics: arithmetic or geometric (default: geometric)

Model Variants:

  • Quantized (models-quantized.tar.gz): ~224MB total (text: 131MB, vision: 92MB, tokenizer: 696KB), faster inference, ~99% quality. Recommended for most use cases.
  • Full (models-full.tar.gz): ~880MB total (text: 522MB, vision: 357MB, tokenizer: 696KB), best quality (fp32 precision). Use for maximum accuracy.

What's Changed

New Contributors

Full Changelog: v20260106...v20260118