Repository navigation
Embeddings and Model Identity
Slackquery uses one logical embedding generation across document indexing and semantic queries. Backend transport can change only when the generated vector space remains provably identical.
| Property | Required value |
|---|---|
| Model | nomic-embed-text:v1.5 |
| Model revision | Operator-provided exact server digest recommended |
| Native dimension | 768 |
| Stored dimension | 512 |
| Document prefix | search_document: |
| Query prefix | search_query: |
| Transformation | first 512 dimensions, then L2 normalization |
| Similarity | cosine |
| Text recipes | Versioned per document kind |
The ordering matters: Slackquery first takes native values [0:512], validates
that all are finite and have nonzero norm, and only then L2-normalizes. Every
stored document vector and every query vector follows the same transformation.
EMBEDDING_BACKEND=pytorch
PYTORCH_EMBEDDING_BASE_URL=http://embedding-host:11435
SLACKQUERY_PYTORCH_EMBEDDING_BATCH_SIZE=1
The server is Ollama API-compatible for /api/embed and /api/tags. Slackquery
also calls /health and requires:
- status
okorhealthy; - expected model/Ollama name;
- a device string beginning with
cuda; - native dimension
768.
EMBEDDING_BACKEND=ollama
OLLAMA_EMBEDDING_BASE_URL=http://ollama-host:11434
Ollama must expose the same model, exact digest, and native dimension through
/api/tags. Do not switch a failed backfill to another transport without
operator approval; both endpoints may share one GPU.
EMBEDDING_BASE_URL takes precedence over the selected backend-specific URL.
Corresponding SLACKQUERY_ aliases remain accepted.
uv run slackquery embedding-statusThis command contacts the selected backend, performs the applicable health and identity checks, and reports backend, URL, model, revision, native/stored dimension, device, and health metadata.
Do not disable verification to force an unrecognized model into an existing generation. Equal model names are insufficient; the digest and transformation contract matter.
The configured generation ID is derived from the model name, full model revision, stored dimension, and recipe version. Durable generation metadata additionally records native dimension, prefixes, normalization, metric, text recipe, and a configuration hash.
Backend name and URL are observability metadata, not generation identity. This is
why switching between the validated PyTorch and Ollama transports does not by
itself invalidate representative message corpus existing vectors.
A change to any vector-defining property requires a separate generation:
- model or digest;
- native/stored dimensional contract;
- document or query prefix;
- truncation order or selected dimensions;
- normalization policy;
- distance metric;
- embedding text recipe.
Never mix vectors from two generations in one artifact.
Projected document text includes workspace, channel, author, and message text.
The embedding client prepends search_document: . Semantic queries are stripped,
length-limited, and sent with search_query: .
The backend request asks for truncation and uses the configured keep-alive value. Slackquery still performs its own structural checks and first-512/L2 transform so the stored contract does not depend on transport behavior alone.
The embedding worker:
- verifies model identity;
- creates or updates generation metadata;
- discovers active projected content without a matching successful vector;
- claims a bounded batch with a lease;
- sends batched
/api/embedrequests; - validates each returned vector;
- checkpoints success or classified retryable/terminal failure;
- repeats until bounded work or available work is exhausted.
Retry behavior uses exponential backoff with jitter and respects Retry-After
for applicable HTTP failures. Expired leases can be reclaimed. Successful vectors
survive process and Dagster run restarts.
To switch from PyTorch to Ollama:
EMBEDDING_BACKEND=ollama uv run slackquery embedding-statusTo switch back:
EMBEDDING_BACKEND=pytorch uv run slackquery embedding-statusBefore persisting the change:
- Confirm the exact full digest, not only the shorthand.
- Confirm native dimension
768and stored dimension512. - Confirm the expected prefixes and normalization remain unchanged.
- Run a semantic smoke query.
- Confirm a no-change embedding run does not enqueue all documents.
Do not re-embed solely because transport changed between these compatible servers. Do re-embed into a new generation if vector identity changed.
The PyTorch CUDA endpoint was measured after warmup with synthetic short document inputs:
| Batch | Elapsed | Throughput |
|---|---|---|
| 1 | 36.3 ms | 27.6 docs/s |
| 32 | 155.6 ms | 205.7 docs/s |
| 128 | 602.6 ms | 212.4 docs/s |
See Testing-and-Benchmarks for the complete benchmark record and its limits.
slackquery wiki
🏠 Overview
🚀 Operate
🔎 Search internals
🔌 Integrate
Sister project