Skip to content

v0.2.0 - Working vector search and `rag ask`

Choose a tag to compare

@hypnagonia hypnagonia released this 06 Sep 08:32
· 6 commits to main since this release
eaad85c

Vector search worked on paper and not in practice. This release makes it actually retrieve, measures how well, and adds a command that answers questions end to end.

Vector search actually retrieves

HybridRetriever was not hybrid. It ran BM25 first, then re-scored only the BM25 candidates, so a chunk BM25 missed could never be retrieved no matter how well it matched semantically. It also ignored rrf_k entirely despite the docs promising RRF. Both arms now run independently and fuse with weighted RRF. (#11)

Alongside that:

  • Embedding is incremental. Every rag index used to re-embed the whole corpus. It now embeds only chunks without a vector and drops vectors for chunks that no longer exist. --force-embed re-embeds everything.
  • Stale vectors are deleted — on document removal and on rebuild. They previously accumulated forever and survived a config change, so Count() > 0 falsely reported that embeddings existed.
  • Dimension is probed from the provider instead of guessed from a hardcoded model table. An unlisted model silently got 1536, then every write failed. Model and dimension are recorded, so a mismatch reports a clear error instead of silently degrading to BM25.
  • rag pack uses embeddings. It was hardcoded to BM25 regardless of config.

Retrieval quality, measured

Built a 14-question ground-truth set and measured instead of guessing (#12):

Config Vector recall@10 MRR
nomic-embed-text @ ~2200-char chunks 2/14 0.054
mxbai-embed-large @ ~540-char chunks 6/14 0.329

Vector recall@200 is 10/14 — the passages are retrieved but ranked 32-141, so the remaining gap needs reranking rather than more embedding tuning. Tuning bm25_weight was measured and does not help (flat at 6/14 across 0.2-0.65).

embedding.include_path is new: prefixing embedded chunks with path:lines helps code but cost 0.12 MRR on prose.

rag ask

The agentic loop lived only in a 1,299-line example that duplicated an OpenAI client and the retriever wiring, so the CLI itself could not answer anything and new features never reached it. It is now rag ask, built on the shared adapters; the example is 172 lines. (#17)

rag ask -d ./books -q "How did Ned Stark die" --fast --hyde

HyDE (--hyde) asks the LLM for a hypothetical answer and searches with that, since questions and answers do not resemble each other in embedding space. One LLM call, cached in the index, so repeats cost zero. On the test corpus it changed a wrong answer into a correct cited one at the same cost.

Every run prints what it used:

   Back-and-forth rounds:  1 of 2 max
   LLM calls:              1
   Total tokens:           3,601  (reported by the API)

Token counts come from the provider's usage field, and the output says whether they were reported or estimated. (#14, #15)

Other fixes

  • rag index /tmp indexed nothing on macOS, where /tmp is a symlink. filepath.Walk lstats its root, so a symlinked root yielded zero files while the CLI's own directory check passed — a silent no-op. (#13)
  • .env is read automatically, searched upward from the target and working directories. Environment variables still win. (#16)
  • Generation is hosted-only. No local LLM provider, in the adapter or the example; a test enforces it.
  • --semantic and --lexical on both query and ask, and they now reject being combined. (#18)
  • Default config had two corrupted globs (***.py, **/node_modulesvendor/**), so any project without a rag.yaml indexed almost nothing.

Known limitations

  • Recall@10 of 6/14 is honest, not good. The answers sit at rank 32-141; a reranking stage over a deep candidate pool is the next real improvement.
  • The agentic loop's context-sufficiency judge is too lenient — it reports "sufficient" on context that yields a wrong answer, so --max-iters 2 often behaves like --fast plus one wasted call.

Note: this tag is v0.2.0 rather than v0.2, so it is valid semver for Go modules.