Skip to content

Embeddings and Model Identity

Thomas Maerz edited this page Oct 4, 2026 · 2 revisions

Embeddings and Model Identity

Slackquery uses one logical embedding generation across document indexing and semantic queries. Backend transport can change only when the generated vector space remains provably identical.

Model contract

Property Required value
Model nomic-embed-text:v1.5
Model revision Operator-provided exact server digest recommended
Native dimension 768
Stored dimension 512
Document prefix search_document:
Query prefix search_query:
Transformation first 512 dimensions, then L2 normalization
Similarity cosine
Text recipes Versioned per document kind

The ordering matters: Slackquery first takes native values [0:512], validates that all are finite and have nonzero norm, and only then L2-normalizes. Every stored document vector and every query vector follows the same transformation.

Backends

PyTorch CUDA, default

EMBEDDING_BACKEND=pytorch
PYTORCH_EMBEDDING_BASE_URL=http://embedding-host:11435
SLACKQUERY_PYTORCH_EMBEDDING_BATCH_SIZE=1

The server is Ollama API-compatible for /api/embed and /api/tags. Slackquery also calls /health and requires:

  • status ok or healthy;
  • expected model/Ollama name;
  • a device string beginning with cuda;
  • native dimension 768.

Ollama transport

EMBEDDING_BACKEND=ollama
OLLAMA_EMBEDDING_BASE_URL=http://ollama-host:11434

Ollama must expose the same model, exact digest, and native dimension through /api/tags. Do not switch a failed backfill to another transport without operator approval; both endpoints may share one GPU.

Common override

EMBEDDING_BASE_URL takes precedence over the selected backend-specific URL. Corresponding SLACKQUERY_ aliases remain accepted.

Verify before use

uv run slackquery embedding-status

This command contacts the selected backend, performs the applicable health and identity checks, and reports backend, URL, model, revision, native/stored dimension, device, and health metadata.

Do not disable verification to force an unrecognized model into an existing generation. Equal model names are insufficient; the digest and transformation contract matter.

Generation identity

The configured generation ID is derived from the model name, full model revision, stored dimension, and recipe version. Durable generation metadata additionally records native dimension, prefixes, normalization, metric, text recipe, and a configuration hash.

Backend name and URL are observability metadata, not generation identity. This is why switching between the validated PyTorch and Ollama transports does not by itself invalidate representative message corpus existing vectors.

A change to any vector-defining property requires a separate generation:

  • model or digest;
  • native/stored dimensional contract;
  • document or query prefix;
  • truncation order or selected dimensions;
  • normalization policy;
  • distance metric;
  • embedding text recipe.

Never mix vectors from two generations in one artifact.

Document and query inputs

Projected document text includes workspace, channel, author, and message text. The embedding client prepends search_document: . Semantic queries are stripped, length-limited, and sent with search_query: .

The backend request asks for truncation and uses the configured keep-alive value. Slackquery still performs its own structural checks and first-512/L2 transform so the stored contract does not depend on transport behavior alone.

Durable worker behavior

The embedding worker:

  1. verifies model identity;
  2. creates or updates generation metadata;
  3. discovers active projected content without a matching successful vector;
  4. claims a bounded batch with a lease;
  5. sends batched /api/embed requests;
  6. validates each returned vector;
  7. checkpoints success or classified retryable/terminal failure;
  8. repeats until bounded work or available work is exhausted.

Retry behavior uses exponential backoff with jitter and respects Retry-After for applicable HTTP failures. Expired leases can be reclaimed. Successful vectors survive process and Dagster run restarts.

Backend switching procedure

To switch from PyTorch to Ollama:

EMBEDDING_BACKEND=ollama uv run slackquery embedding-status

To switch back:

EMBEDDING_BACKEND=pytorch uv run slackquery embedding-status

Before persisting the change:

  1. Confirm the exact full digest, not only the shorthand.
  2. Confirm native dimension 768 and stored dimension 512.
  3. Confirm the expected prefixes and normalization remain unchanged.
  4. Run a semantic smoke query.
  5. Confirm a no-change embedding run does not enqueue all documents.

Do not re-embed solely because transport changed between these compatible servers. Do re-embed into a new generation if vector identity changed.

Measured throughput

The PyTorch CUDA endpoint was measured after warmup with synthetic short document inputs:

Batch Elapsed Throughput
1 36.3 ms 27.6 docs/s
32 155.6 ms 205.7 docs/s
128 602.6 ms 212.4 docs/s

See Testing-and-Benchmarks for the complete benchmark record and its limits.

Clone this wiki locally