Skip to content

Architecture

Thomas Maerz edited this page Oct 4, 2026 · 2 revisions

Architecture

Slackquery separates canonical ingestion, mutable enrichment state, immutable serving artifacts, and read-only retrieval. This prevents the MCP service from contending with pipeline writers and gives publication a clear validation and rollback boundary.

System topology

flowchart TB
    subgraph Canonical["Slackpipe-owned canonical layer"]
      CDB["Canonical Slack DuckDB"]
      FILES["Attachment tree"]
    end

    subgraph Pipeline["Slackquery pipeline / Dagster"]
      P["document_projection"]
      ES["message_embeddings"]
      BA["artifact_candidate"]
      CK["candidate_integrity\nblocking check"]
      PA["published_artifact"]
      SDB["Slackquery state DuckDB"]
    end

    subgraph Models["Local embedding transport"]
      PT["PyTorch-compatible CUDA endpoint"]
      OL["Optional Ollama transport"]
    end

    subgraph Serving["Immutable serving layer"]
      ART["Immutable search artifacts"]
      CUR["current.duckdb"]
      MCP["MCP :8181/mcp"]
    end

    CDB -->|"ATTACH READ_ONLY"| P
    FILES -->|"read-only safe extraction"| P
    P --> SDB
    SDB --> ES
    ES <--> PT
    ES -.->|"backend switch"| OL
    ES --> SDB
    SDB --> BA
    BA --> ART
    ART --> CK
    CK --> PA
    PA --> CUR
    CUR -->|"fresh read-only connections"| MCP
Loading

Data flow

1. Canonical projection

Projection attaches the canonical database with DuckDB READ_ONLY, verifies its schema, and materializes deterministic message, thread-context, and attachment chunk documents. Attachment paths must resolve inside configured workspace roots; traversal and symlink escapes are rejected.

The embedding text recipe is message-v1:

workspace: <workspace slug>
channel: <channel name or ID>
author: <best available display name or ID>
message: <message text>

Deleted messages and messages with empty canonical search text are inactive and are excluded from new artifacts.

2. Durable embedding state

The state database stores:

  • projected documents and source watermarks;
  • embedding-generation identity and configuration;
  • per-document embedding state, attempts, leases, errors, and vectors;
  • artifact build history, paths, checksums, statuses, and Dagster run IDs.

Embedding rows are content-addressed by document, source version, text hash, and generation. Completed vectors survive worker and process restarts.

3. Immutable artifact build

Build requires a successful vector for every active document in the selected generation. It creates a temporary DuckDB containing:

  • search_documents: active message metadata and search/display text;
  • document_vectors: one FLOAT[512] vector per document;
  • artifact_metadata: build, schema, generation, watermark, counts, timestamp, and FTS configuration.

DuckDB FTS indexes text_lexical with no stemmer, no stopwords, lowercase matching, and accent stripping. The build is checkpointed, validated through a fresh read-only connection, renamed, made read-only, checksummed, and accompanied by a JSON manifest.

4. Atomic publication

Publication validates the artifact and manifest again, then atomically replaces /srv/slackquery/artifacts/current.duckdb with a relative symlink. Directory fsync makes the selector replacement durable. The default retention policy keeps three artifacts and cannot be configured below two.

5. Read-only serving

The MCP process resolves current.duckdb and opens a fresh read-only DuckDB connection for retrieval. It does not write artifact or pipeline state and does not expose arbitrary SQL or filesystem access.

Component boundaries

Component May read May write
Slackpipe Its canonical sources and database Canonical Slackpipe database
Slackquery projection/embedding Canonical DB, Slackquery state Slackquery state only
Slackquery artifact builder Slackquery state New artifact, manifest, build state
Slackquery publisher Validated artifacts/state current.duckdb, build status, retention deletes
MCP service Published artifact, embedding query endpoint No database writes

Serving schema and API records

Search hits carry document/workspace/channel/author IDs and names, timestamp, thread IDs, text, permalink, component ranks/scores, and fused score. Search responses include the query, mode, artifact ID, result list, and optional opaque cursor. This makes result provenance explicit and lets clients detect which artifact answered a request.

Failure isolation

  • Model or worker failure leaves existing vectors and published artifact intact.
  • Candidate build failure never mutates the current artifact.
  • Integrity-check failure blocks the Dagster publication asset.
  • Publication is a selector change rather than an in-place database mutation.
  • A broken or absent selector makes readiness fail instead of serving partial state.
  • Retained immutable artifacts support rollback after checksum validation.

Current scale

The validated deployment serves representative message corpus active message documents and vectors. The observed artifact is /srv/slackquery/artifacts/search-<build-id>.duckdb. Exact cosine remains deliberately simpler and sufficiently fast at this scale; see Testing-and-Benchmarks and Design-Decisions.

Clone this wiki locally