Skip to content

Retrieval and RRF

Thomas Maerz edited this page Oct 4, 2026 · 2 revisions

Retrieval and RRF

Slackquery supports lexical, semantic, and hybrid retrieval over one immutable artifact. Results can be messages or file chunks; thread-context documents add ranking context and map back to message evidence.

Lexical retrieval

Lexical mode uses DuckDB FTS/BM25 over text_lexical.

The artifact FTS configuration is:

Option Value
Stemmer none
Stopwords none
Lowercase enabled
Accent stripping enabled

Query text is capped, tokenized with a conservative allowlist, and limited to 64 tokens before it is safely quoted for the generated FTS macro. Metadata filter values remain parameterized. Lexical mode is best for exact errors, identifiers, URLs, filenames, code, acronyms, names, and distinctive phrases.

Semantic retrieval

Semantic mode:

  1. prepends search_query: to the query;
  2. calls the validated local embedding endpoint;
  3. takes native dimensions 0:512;
  4. L2-normalizes the vector;
  5. computes exact DuckDB array_cosine_similarity against filtered stored vectors;
  6. ranks by descending cosine score.

This is an exact scan, not ANN/HNSW. It avoids approximate-index recall risk and experimental persisted-index behavior until measured SLO evidence justifies it.

Hybrid retrieval

Hybrid mode classifies queries as exact, conceptual, or mixed, then retrieves lexical, semantic, and contextual candidates. It does not add BM25 and cosine values because their scales and query-dependent distributions are incompatible.

For document d, weighted reciprocal rank fusion is:

RRF(d) = lexical contribution + semantic contribution + context contribution

source contribution = source_weight / (60 + source_rank(d))

Defaults:

Parameter Value
RRF k 60
Candidate window per retriever 100
Lexical weight 1.0
Semantic weight 1.0

The candidate union is ordered by fused score and a stable document-ID tie-break, then joined back to message metadata. Component rank and score are returned with the fused score.

RRF is a robust baseline, not a completed relevance finding. A human-judged query set still must compare lexical-only, semantic-only, default RRF, and a small parameter grid. Do not claim that hybrid relevance has been proven superior.

Filters

search_slack supports:

  • workspace_ids;
  • channel_ids;
  • author_ids;
  • start_ts_us and end_ts_us inclusive bounds;
  • message document kind internally;
  • limit and opaque cursor.

Filters are applied to each retriever before fusion. Time values are Unix epoch microseconds. The default maximum is 100 values per filter list and 50 returned results per call.

Pagination

The cursor contains an encoded artifact ID, a fingerprint of query/mode/filters, and an offset. It is intentionally opaque and validated on reuse.

A cursor becomes invalid when:

  • a different artifact is published;
  • query text changes;
  • mode changes;
  • any filter changes;
  • it is malformed or its offset is out of range.

Restart pagination without the cursor after an hourly artifact change.

Message and thread expansion

Retrieval ranks atomic messages. This preserves precise citations and avoids duplicating or truncating whole-thread text in every vector.

  • get_slack_message(document_id, context_before, context_after) fetches one exact message and bounded chronological same-channel neighbors.
  • get_slack_thread(thread_id, limit) fetches the matching thread chronologically.

Context should be requested after identifying a promising hit. Same-channel neighbors are not necessarily part of the same discussion; use thread expansion for thread claims.

Interpreting results

  • BM25, cosine, and RRF scores are ranking signals, not probabilities.
  • A high semantic score is not proof of factual equivalence.
  • A zero-result response is not proof of absence. Filters, archive coverage, deletion, empty text, naming, and spelling can hide relevant messages.
  • Cite workspace, channel, author, timestamp, document_id, and permalink when available.

Measured latency

Eight representative queries, three iterations each, produced:

Mode Requests p50 p95 Maximum
Lexical BM25 24 70.62 ms 79.10 ms 84.05 ms
Semantic exact cosine 24 219.67 ms 259.97 ms 301.60 ms
Hybrid RRF 24 266.94 ms 309.69 ms 329.19 ms

These are latency measurements, not human relevance results. See Testing-and-Benchmarks.

Clone this wiki locally