Skip to content
Thomas Maerz edited this page Oct 4, 2026 · 3 revisions

slackquery

Slackquery provides production-oriented hybrid search over a canonical Slack archive. It deterministically projects active messages, bounded thread contexts, and safe attachment chunks; creates content-addressed embeddings with a verified local model; builds immutable DuckDB artifacts; and serves structured retrieval tools through MCP Streamable HTTP.

Slackquery is the search sister project to Slackpipe, which owns multi-workspace Slack extraction and canonical DuckDB production. See the Slackpipe technical documentation for the upstream data contract.

Example deployment

Component Address or status
Repository https://github.com/thomasmaerz/slackquery
MCP endpoint http://mcp-host:8181/mcp
Liveness http://mcp-host:8181/healthz
Readiness http://mcp-host:8181/readyz
Metrics http://mcp-host:8181/metrics
Dagster http://dagster-host:3000
Dagster code location slackquery
Reconciliation schedule slackquery_hourly_reconciliation

Replace example hosts with local values. Use current health, Dagster run, and asset status when diagnosing a deployment.

Capabilities

  • DuckDB FTS/BM25 lexical retrieval.
  • Exact cosine semantic retrieval over normalized 512-dimensional vectors.
  • Query-aware exact/conceptual/mixed routing and weighted reciprocal rank fusion.
  • Exact phrase/token boosts and deterministic thread/channel diversity.
  • Bounded thread contexts and safe text/code/CSV/PDF/DOCX/PPTX chunks.
  • Workspace, channel, author, and timestamp filters.
  • Message lookup, bounded channel context, and chronological thread expansion.
  • Immutable, checksummed artifacts selected by an atomic symlink.
  • Strict embedding model/digest/dimension validation.
  • Hourly Dagster projection, embedding reconciliation, build, integrity check, and publication.
  • Prometheus metrics, retention, rollback, and state backup/restore.

Architecture at a glance

flowchart LR
    A["Slackpipe canonical DuckDB\nread-only"] --> B["Projection"]
    B --> C["Slackquery state DuckDB"]
    C <--> D["PyTorch CUDA / Ollama\nembedding backend"]
    C --> E["Immutable DuckDB build\nFTS + vectors"]
    E --> F["current.duckdb"]
    F --> G["MCP Streamable HTTP"]
    G --> H["MCP clients"]
    I["Dagster hourly reconciliation"] --> B
    I --> D
    I --> E
    I --> F
Loading

See Architecture for boundaries and data flow.

MCP interface

Slackquery exposes four read-only tools:

Tool Use
search_slack Search in lexical, semantic, or hybrid mode.
get_slack_message Fetch one message and optional bounded same-channel context.
get_slack_thread Fetch a thread chronologically.
list_slack_scopes Discover stable workspace/channel IDs and archive coverage.

It also exposes slackquery://guide. Start with hybrid retrieval, use lexical for exact strings and semantic for paraphrases, and expand only promising messages or threads. See MCP-Client-Setup.

Current status: Bronze implemented, promotion incomplete

Bronze, Silver, and Gold describe evidence-gated capability levels.

Bronze

The main Bronze implementation is operational: message-only projection, one verified embedding generation, immutable artifacts, FTS, exact cosine, equal RRF, four MCP tools, integrity checks, atomic publication, rollback retention, and an hourly schedule.

Bronze is not fully promoted because the required human-judged relevance benchmark has not been completed. Latency and throughput have been measured, but those measurements do not establish relevance quality.

Silver

Some foundations exist, notably content-addressed incremental embeddings, checkpointing, and scheduled repair. Evaluated contextual thread representations, file-chunk retrieval, explicit quality regression gates, and complete generation migration procedures remain outstanding.

Gold

ANN, cross-encoder reranking, learned routing/fusion, blue/green serving, full alerting/SLO ownership, and rehearsed disaster recovery remain future work. ANN is intentionally deferred while exact scan meets the current interactive LAN latency target.

See Design-Decisions and Testing-and-Benchmarks.

Production paths

Purpose Path
Canonical source, read-only /path/to/slackpipe.duckdb
Durable enrichment state /srv/slackquery/state/slackquery.duckdb
Artifacts and manifests /srv/slackquery/artifacts
Published selector /srv/slackquery/artifacts/current.duckdb
DuckDB extension cache /srv/slackquery/extensions

Slackpipe owns the canonical database. Slackquery owns its state, artifacts, and extension cache. The MCP service opens only the selected artifact read-only.

Read next

The repository README and maintained docs are at https://github.com/thomasmaerz/slackquery.

Clone this wiki locally