Skip to content

Troubleshooting

Thomas Maerz edited this page Oct 4, 2026 · 2 revisions

Troubleshooting

First-response checklist

  1. Check MCP liveness and readiness.
  2. Check Dagster instance/code-location health, recent schedule ticks, and latest slackquery_reconcile run.
  3. Run slackquery embedding-status for semantic/hybrid or embedding failures.
  4. Identify the exact current artifact and validate its checksum if artifact integrity is in question.
  5. Preserve the current known-good artifact; do not mutate it in place.
curl -fsS http://mcp-host:8181/healthz
curl -fsS http://mcp-host:8181/readyz
uv run slackquery embedding-status

MCP process unavailable

Symptoms

  • connection refused;
  • timeout;
  • /healthz unavailable.

Checks

  • service/container status and restart loop;
  • configured host/port and host port mapping (8181 to container 8080);
  • loopback versus LAN bind address;
  • firewall/routing;
  • read-only-root mount requirements;
  • service logs for startup configuration errors.

Readiness returns 503

Readiness opens the selected artifact and reads metadata.

Check:

  • the configured current.duckdb selector exists;
  • it is not a broken symlink;
  • its target exists inside the artifact directory;
  • the service user can read the target and extension directory;
  • the artifact validates with checksum.

Recover by publishing a validated known-good artifact. Do not copy a mutable database over current.duckdb.

401 Unauthorized

SLACKQUERY_BEARER_TOKEN is configured or the supplied token is wrong. MCP requests require:

Authorization: Bearer <token>

Health endpoints remain public. Check client secret interpolation and whitespace without printing or logging the token.

Model verification fails

The required identity is:

  • model nomic-embed-text:v1.5;
  • digest operator-pinned-model-revision;
  • native dimension 768;
  • stored dimension 512 after first-512 truncation and L2 normalization.

Both backends must report matching /api/tags. PyTorch additionally must report healthy CUDA operation through /health.

Do not change the configured digest or disable checks to accept an unknown model. A legitimate model change requires a new generation.

PyTorch unavailable

Check http://embedding-host:11435/health from the runtime network. If necessary, switch to:

EMBEDDING_BACKEND=ollama

The fallback at http://ollama-host:11434 is valid only if it reports the same model digest and dimension. Run embedding-status after switching.

Semantic/hybrid fails while lexical works

Lexical mode does not call the query embedding backend. Semantic and hybrid modes do. This pattern usually indicates backend network, health, identity, timeout, or response validation failure. Diagnose the backend rather than treating lexical results as semantically equivalent.

embedding coverage incomplete: X/Y

Build found active projected documents without successful current-generation vectors.

  1. Run projection to ensure state is current.
  2. Repeat bounded or full embedding work.
  3. Inspect retryable/terminal errors in state and backend logs.
  4. Correct the root cause and retry.
  5. Build only when active document and successful vector counts agree.

Do not manually mark failed rows successful or bypass the build guard.

Embedding appears to restart from zero

Check whether model digest, dimension, prefix, normalization, text recipe, or generation configuration changed. Switching only between the validated PyTorch and Ollama transports should not create a new generation. Also check whether the state database path or mount changed, causing an empty state DB to be created.

FTS extension failure

Ensure /srv/slackquery/extensions:

  • exists;
  • is writable during extension installation/build;
  • is mounted into both build/Dagster and MCP runtime environments;
  • is configured as DuckDB's extension directory.

The container root is read-only, so installing into $HOME/.duckdb is not a working fallback.

Candidate or integrity check fails

Inspect the Dagster step and check metadata. Validate:

  • complete active document/vector coverage;
  • no duplicate document IDs;
  • one-to-one document/vector key sets;
  • vector length exactly 512;
  • normalized, finite, nonzero vectors;
  • artifact metadata counts;
  • manifest SHA-256;
  • disk space and artifact-directory permissions.

The currently published artifact remains untouched. Do not publish manually around the blocking check.

Cursor rejected

The cursor is bound to artifact ID plus query, mode, and filters. Hourly publication or any request change invalidates it. Restart pagination without the cursor.

Missing or surprising search result

  • Use list_slack_scopes to verify stable IDs and first/last archive timestamps.
  • Remove overly narrow workspace/channel/author/time filters.
  • Use lexical mode for exact strings and semantic mode for paraphrases.
  • Check spelling, deletion state, and whether message search text is empty.
  • Expand promising threads before judging context.

No result does not prove absence from Slack or even from all archived data.

Schedule is RUNNING, but data is stale

RUNNING only means enabled. Check:

  • recent tick status and launched run IDs;
  • skipped or failed ticks;
  • latest run status and failed step;
  • code-location load health;
  • projection watermark and materialization timestamps;
  • embedding failure counts;
  • candidate check and publication status;
  • MCP artifact_id after the expected run.

Rollback

Validate and publish a retained artifact:

uv run slackquery validate --checksum \
  /srv/slackquery/artifacts/search-<known-good-build>.duckdb
uv run slackquery publish \
  /srv/slackquery/artifacts/search-<known-good-build>.duckdb

Then verify readiness and smoke-test lexical, semantic, and hybrid retrieval. If the scheduled job would recreate the bad state, pause/remediate orchestration as part of incident handling.

See Deployment-and-Operations and Dagster-Orchestration.

Clone this wiki locally