Repository navigation
v0.2.0 - Working vector search and `rag ask`
Vector search worked on paper and not in practice. This release makes it actually retrieve, measures how well, and adds a command that answers questions end to end.
Vector search actually retrieves
HybridRetriever was not hybrid. It ran BM25 first, then re-scored only the BM25 candidates, so a chunk BM25 missed could never be retrieved no matter how well it matched semantically. It also ignored rrf_k entirely despite the docs promising RRF. Both arms now run independently and fuse with weighted RRF. (#11)
Alongside that:
- Embedding is incremental. Every
rag indexused to re-embed the whole corpus. It now embeds only chunks without a vector and drops vectors for chunks that no longer exist.--force-embedre-embeds everything. - Stale vectors are deleted — on document removal and on rebuild. They previously accumulated forever and survived a config change, so
Count() > 0falsely reported that embeddings existed. - Dimension is probed from the provider instead of guessed from a hardcoded model table. An unlisted model silently got 1536, then every write failed. Model and dimension are recorded, so a mismatch reports a clear error instead of silently degrading to BM25.
rag packuses embeddings. It was hardcoded to BM25 regardless of config.
Retrieval quality, measured
Built a 14-question ground-truth set and measured instead of guessing (#12):
| Config | Vector recall@10 | MRR |
|---|---|---|
nomic-embed-text @ ~2200-char chunks |
2/14 | 0.054 |
mxbai-embed-large @ ~540-char chunks |
6/14 | 0.329 |
Vector recall@200 is 10/14 — the passages are retrieved but ranked 32-141, so the remaining gap needs reranking rather than more embedding tuning. Tuning bm25_weight was measured and does not help (flat at 6/14 across 0.2-0.65).
embedding.include_path is new: prefixing embedded chunks with path:lines helps code but cost 0.12 MRR on prose.
rag ask
The agentic loop lived only in a 1,299-line example that duplicated an OpenAI client and the retriever wiring, so the CLI itself could not answer anything and new features never reached it. It is now rag ask, built on the shared adapters; the example is 172 lines. (#17)
rag ask -d ./books -q "How did Ned Stark die" --fast --hydeHyDE (--hyde) asks the LLM for a hypothetical answer and searches with that, since questions and answers do not resemble each other in embedding space. One LLM call, cached in the index, so repeats cost zero. On the test corpus it changed a wrong answer into a correct cited one at the same cost.
Every run prints what it used:
Back-and-forth rounds: 1 of 2 max
LLM calls: 1
Total tokens: 3,601 (reported by the API)
Token counts come from the provider's usage field, and the output says whether they were reported or estimated. (#14, #15)
Other fixes
rag index /tmpindexed nothing on macOS, where/tmpis a symlink.filepath.Walklstats its root, so a symlinked root yielded zero files while the CLI's own directory check passed — a silent no-op. (#13).envis read automatically, searched upward from the target and working directories. Environment variables still win. (#16)- Generation is hosted-only. No local LLM provider, in the adapter or the example; a test enforces it.
--semanticand--lexicalon bothqueryandask, and they now reject being combined. (#18)- Default config had two corrupted globs (
***.py,**/node_modulesvendor/**), so any project without arag.yamlindexed almost nothing.
Known limitations
- Recall@10 of 6/14 is honest, not good. The answers sit at rank 32-141; a reranking stage over a deep candidate pool is the next real improvement.
- The agentic loop's context-sufficiency judge is too lenient — it reports "sufficient" on context that yields a wrong answer, so
--max-iters 2often behaves like--fastplus one wasted call.
Note: this tag is v0.2.0 rather than v0.2, so it is valid semver for Go modules.