Skip to content

Testing and Benchmarks

Thomas Maerz edited this page Oct 4, 2026 · 2 revisions

Testing and Benchmarks

Repository checks

Run from the repository root:

uv run pytest
uv run ruff check .
uv run mypy src/slackquery

The test suite covers:

  • canonical projection and tombstoning;
  • thread contexts, safe attachment extraction, and deterministic chunking;
  • content-addressed, resumable embedding behavior;
  • strict model identity and vector validation;
  • immutable build, validation, publication, and retention behavior;
  • lexical, semantic, and hybrid retrieval;
  • route-aware weighted RRF, exact boosts, and diversity;
  • FTS query sanitization and parameterized filters;
  • cursor binding and rejection;
  • MCP tool/resource and health/readiness contracts;
  • Dagster definitions and blocking integrity checks;
  • container/deployment contracts.
  • retention, rollback, backup/restore, and Gold validation.

Measurement context

The maintained benchmark record predates Gold thread/file projection. It used a representative message-only corpus with normalized 512-dimensional vectors and exact cosine scan.

The benchmark source exists in the repository at docs/benchmarks.md.

Embedding endpoint

A local PyTorch CUDA server was measured with synthetic short document inputs after warmup:

Batch Elapsed Throughput
1 36.3 ms 27.6 docs/s
32 155.6 ms 205.7 docs/s
128 602.6 ms 212.4 docs/s

These historical throughput measurements are not safe production batch defaults for every GPU. Use SLACKQUERY_PYTORCH_EMBEDDING_BATCH_SIZE and validate the target device. Compatible transports can share vectors only when model revision, prefixes, truncation, and normalization are identical.

Retrieval latency

Eight representative queries, three iterations each:

Mode Requests p50 p95 Maximum
Lexical BM25 24 70.62 ms 79.10 ms 84.05 ms
Semantic exact cosine 24 219.67 ms 259.97 ms 301.60 ms
Hybrid RRF 24 266.94 ms 309.69 ms 329.19 ms

Exact scan meets the current interactive LAN target, so persisted experimental HNSW is not currently justified.

Artifact build

Building and indexing the representative message corpus-document immutable DuckDB artifact took 6.75 seconds. Artifact checksum validation and a fresh read-only open passed before publication.

The observed published artifact is:

/srv/slackquery/artifacts/search-<build-id>.duckdb

What these numbers do not prove

These measurements characterize throughput and latency. They do not establish retrieval relevance, answer correctness, or user satisfaction. In particular:

  • the human-judged relevance benchmark has not been completed;
  • RRF has not yet been proven no worse than the best single retriever on a judged set;
  • no nDCG, MRR, Recall@k, or success@k result should be inferred;
  • Bronze promotion remains incomplete on the relevance gate.

Required relevance evaluation

The design calls for a versioned set of at least 100–200 real information needs covering:

  • exact errors and exception text;
  • tickets, commits, URLs, filenames, products, people, channels, and acronyms;
  • paraphrased how/why/decision questions;
  • incident chronology and thread-specific questions;
  • rare terms, misspellings, and no-answer cases;
  • eventually file-to-message and message-to-file discovery.

Candidates should be pooled from BM25, semantic retrieval, multiple RRF settings, and future retrievers. Knowledgeable reviewers should grade answer-bearing messages/threads. Planned metrics include nDCG@10, MRR@10, Recall@10/50/100, success@k, zero-result and false-positive rates, and per-slice latency/quality.

Default RRF (k=60, window 100, equal weights) is a baseline to test, not an unchangeable conclusion.

Benchmark hygiene

  • Record corpus count, artifact ID, generation ID, backend, query set version, concurrency, warmup, and iterations.
  • Separate embedding-server time from DuckDB retrieval time where possible.
  • Preserve raw observations and report p50/p95/p99 rather than averages alone.
  • Compare ANN recall to exact top-k before introducing approximate retrieval.
  • Do not tune and evaluate on the same judged queries without documenting the split.
  • Treat production Slack content and judgments as sensitive data.

See Retrieval-and-RRF and Design-Decisions.

Clone this wiki locally