Repository navigation
Testing and Benchmarks
Run from the repository root:
uv run pytest
uv run ruff check .
uv run mypy src/slackqueryThe test suite covers:
- canonical projection and tombstoning;
- thread contexts, safe attachment extraction, and deterministic chunking;
- content-addressed, resumable embedding behavior;
- strict model identity and vector validation;
- immutable build, validation, publication, and retention behavior;
- lexical, semantic, and hybrid retrieval;
- route-aware weighted RRF, exact boosts, and diversity;
- FTS query sanitization and parameterized filters;
- cursor binding and rejection;
- MCP tool/resource and health/readiness contracts;
- Dagster definitions and blocking integrity checks;
- container/deployment contracts.
- retention, rollback, backup/restore, and Gold validation.
The maintained benchmark record predates Gold thread/file projection. It used a representative message-only corpus with normalized 512-dimensional vectors and exact cosine scan.
The benchmark source exists in the repository at
docs/benchmarks.md.
A local PyTorch CUDA server was measured with synthetic short document inputs after warmup:
| Batch | Elapsed | Throughput |
|---|---|---|
| 1 | 36.3 ms | 27.6 docs/s |
| 32 | 155.6 ms | 205.7 docs/s |
| 128 | 602.6 ms | 212.4 docs/s |
These historical throughput measurements are not safe production batch defaults
for every GPU. Use SLACKQUERY_PYTORCH_EMBEDDING_BATCH_SIZE and validate the
target device. Compatible transports can share vectors only when model revision,
prefixes, truncation, and normalization are identical.
Eight representative queries, three iterations each:
| Mode | Requests | p50 | p95 | Maximum |
|---|---|---|---|---|
| Lexical BM25 | 24 | 70.62 ms | 79.10 ms | 84.05 ms |
| Semantic exact cosine | 24 | 219.67 ms | 259.97 ms | 301.60 ms |
| Hybrid RRF | 24 | 266.94 ms | 309.69 ms | 329.19 ms |
Exact scan meets the current interactive LAN target, so persisted experimental HNSW is not currently justified.
Building and indexing the representative message corpus-document immutable DuckDB artifact took
6.75 seconds. Artifact checksum validation and a fresh read-only open passed
before publication.
The observed published artifact is:
/srv/slackquery/artifacts/search-<build-id>.duckdb
These measurements characterize throughput and latency. They do not establish retrieval relevance, answer correctness, or user satisfaction. In particular:
- the human-judged relevance benchmark has not been completed;
- RRF has not yet been proven no worse than the best single retriever on a judged set;
- no nDCG, MRR, Recall@k, or success@k result should be inferred;
- Bronze promotion remains incomplete on the relevance gate.
The design calls for a versioned set of at least 100–200 real information needs covering:
- exact errors and exception text;
- tickets, commits, URLs, filenames, products, people, channels, and acronyms;
- paraphrased how/why/decision questions;
- incident chronology and thread-specific questions;
- rare terms, misspellings, and no-answer cases;
- eventually file-to-message and message-to-file discovery.
Candidates should be pooled from BM25, semantic retrieval, multiple RRF settings, and future retrievers. Knowledgeable reviewers should grade answer-bearing messages/threads. Planned metrics include nDCG@10, MRR@10, Recall@10/50/100, success@k, zero-result and false-positive rates, and per-slice latency/quality.
Default RRF (k=60, window 100, equal weights) is a baseline to test, not an
unchangeable conclusion.
- Record corpus count, artifact ID, generation ID, backend, query set version, concurrency, warmup, and iterations.
- Separate embedding-server time from DuckDB retrieval time where possible.
- Preserve raw observations and report p50/p95/p99 rather than averages alone.
- Compare ANN recall to exact top-k before introducing approximate retrieval.
- Do not tune and evaluate on the same judged queries without documenting the split.
- Treat production Slack content and judgments as sensitive data.
See Retrieval-and-RRF and Design-Decisions.
slackquery wiki
🏠 Overview
🚀 Operate
🔎 Search internals
🔌 Integrate
Sister project