Repository navigation
Dagster Orchestration
Slackquery exposes Dagster definitions from:
slackquery.definitions:definitions
flowchart LR
P["search/document_projection"] --> E["search/message_embeddings"]
E --> C["search/artifact_candidate"]
C --> K["candidate_integrity\nblocking asset check"]
K --> U["search/published_artifact"]
Projects the canonical DuckDB archive into Slackquery state. Materialization metadata reports projection, active, tombstoned, and source-watermark summaries. The canonical database is attached read-only.
Runs the resumable worker against active projected rows. Metadata reports claimed, succeeded, retryable-failed, and terminal-failed work. A healthy no-change run may claim no items because successful vectors are durable.
Builds a new immutable artifact for the selected projection watermark and vector generation. Build lineage is stored in durable state so downstream steps choose the candidate produced by the same orchestration run without relying on a published run identifier in documentation.
This blocking asset check opens the candidate through a fresh read-only connection and verifies:
- schema and required metadata;
- active document/vector coverage;
- unique and matching document keys;
- vector shape, finite values, and normalization;
- manifest checksum.
Publication cannot proceed when the check fails.
Selects the validated candidate for the current run and publishes it by atomically
replacing current.duckdb. Materialization metadata records the selected artifact
without mutating it.
| Definition | Value |
|---|---|
| Asset job | slackquery_reconcile |
| Schedule | slackquery_hourly_reconciliation |
| Cron | 0 * * * * |
The schedule performs periodic reconciliation rather than depending on an upstream event. Projection and embedding are idempotent and content-addressed, so a no-change run can safely validate and publish without re-embedding unchanged content.
Enable and set the schedule timezone according to deployment policy. A schedule
shown as RUNNING is enabled; that status alone does not prove successful ticks,
fresh assets, a passing candidate, or a healthy MCP service.
For a complete health check:
- Verify the Dagster instance, daemons, and code location.
- Inspect recent schedule ticks for failure, skips, and launched runs.
- Inspect the latest
slackquery_reconcilestatus and failed step logs. - Compare asset materialization metadata and vector coverage.
- Confirm
candidate_integritypassed. - Verify MCP
/readyzafter publication and run a search smoke test.
Check the canonical file mount, schema compatibility, read permissions, and database readability. Do not grant canonical write access as a workaround.
Check selected backend, network access, model revision, dimensions, endpoint health, and per-item failure classes. Existing successful vectors remain durable.
Common causes are incomplete vector coverage, unavailable FTS extension, disk space, or permissions. A failed candidate does not affect the published artifact.
Do not publish around the check. Investigate metadata counts, key mismatch, vector validation, and manifest checksum while the known-good artifact continues serving.
Check artifact-directory ownership, selector permissions, checksum, and disk state. Publication should accept candidates only from the configured artifact directory.
Do not run a second manual pipeline writer while scheduled reconciliation is active. DuckDB state is operated as a single-writer store. Read-only MCP traffic remains isolated because it opens the immutable artifact.
See Deployment and Operations and Troubleshooting.
slackquery wiki
🏠 Overview
🚀 Operate
🔎 Search internals
🔌 Integrate
Sister project