Repository navigation
Releases: hypnagonia/rag
Release list
v0.3.0 - Smaller index, 20x faster vector queries, measurable retrieval
Two changes since v0.2.0: the index got much smaller and vector queries got much faster, and retrieval quality is now measurable by anyone rather than by scripts that lived on my machine.
Packed binary vectors (#20)
Measured the index bucket by bucket before touching it. Vectors were 80.6% of it, stored as JSON arrays of floats:
| bucket | keys | MB | share |
|---|---|---|---|
| vectors | 20,340 | 248.0 | 80.6% |
| terms | 13,273 | 37.5 | 12.2% |
| blobs | 20,340 | 11.5 | 3.7% |
| chunks | 20,340 | 10.5 | 3.4% |
Vectors are now packed little-endian float32 with a 5-byte header and an optional JSON metadata suffix. The encoding is lossless, so results are unchanged.
On a 20,340-vector index:
| before | after | |
|---|---|---|
| file size | 439 MB | 192.7 MB (−66.7%) |
| semantic query | 2.58s | 0.13s (~20x) |
| top-3 results | 0.77 / 0.75 / 0.74 | identical |
The latency win was the surprise: JSON decoding on every process start dominated, not the similarity search. Brute-force scan over 20k vectors is ~0.1s.
rag compact is new. It rewrites legacy JSON records in place and reclaims free pages - no re-embedding, no quality change. Legacy records stay readable, so existing indexes keep working untouched:
$ rag compact -d ./books
Rewrote 20340 vector records as packed binary
Index: 579.3 MB -> 192.7 MB (66.7% smaller)
Note the file grows during the rewrite before compaction reclaims it, so leave headroom.
rageval (#21)
Every recall and MRR figure quoted in the README came from throwaway scripts in a scratch directory, which made the measurement rule in CLAUDE.md unfollowable by anyone else. That harness is now a first-class Go command:
go run ./cmd/rageval -corpus /path/to/indexed -k 1014 questions, recall@10, MMR disabled
MODE RECALL MRR RANKS
lexical 6/14 0.201 [- 4 - - - - - - 5 9 1 - 4 1]
semantic 6/14 0.329 [- 1 - - - - - - 1 - 1 2 9 1]
hybrid 6/14 0.310 [- 1 - - - - - - 2 - 1 3 2 1]
Each question is paired with an anchor phrase that must appear in a correctly retrieved passage, so it measures whether the right passage was found rather than whether an answer looked plausible. -modes hyde includes HyDE, -v prints per-question ranks. The bundled questions.json covers A Song of Ice and Fire and is meant to be replaced per corpus.
It has already earned its place: it falsified a plausible-sounding local reranking idea in four minutes, before any of it shipped.
Known limitations
Unchanged from v0.2.0 and worth restating:
- recall@10 is 6/14. recall@200 is 10/14, so the passages are retrieved but ranked 32-141. Closing that needs a cross-encoder reranker, which means genuine query-passage interaction and a second API call. A bi-encoder cannot do it, however the input is sliced.
- The agentic loop's context-sufficiency judge is too lenient, so
--max-iters 2often costs a call without changing the answer. - The index is still ~19x the source text.
terms(JSON postings),chunks(a redundant token array) andblobs(uncompressed text) are ~48 MB of lossless savings not yet taken. - All quality numbers come from 14 questions on one prose corpus. The tool is code-first - AST chunking, symbol extraction, path boosting - and none of that is measured.
embedding.include_pathdefaults to true on the assumption that paths help code, tested only on prose, where it hurt.
v0.2.0 - Working vector search and `rag ask`
Vector search worked on paper and not in practice. This release makes it actually retrieve, measures how well, and adds a command that answers questions end to end.
Vector search actually retrieves
HybridRetriever was not hybrid. It ran BM25 first, then re-scored only the BM25 candidates, so a chunk BM25 missed could never be retrieved no matter how well it matched semantically. It also ignored rrf_k entirely despite the docs promising RRF. Both arms now run independently and fuse with weighted RRF. (#11)
Alongside that:
- Embedding is incremental. Every
rag indexused to re-embed the whole corpus. It now embeds only chunks without a vector and drops vectors for chunks that no longer exist.--force-embedre-embeds everything. - Stale vectors are deleted — on document removal and on rebuild. They previously accumulated forever and survived a config change, so
Count() > 0falsely reported that embeddings existed. - Dimension is probed from the provider instead of guessed from a hardcoded model table. An unlisted model silently got 1536, then every write failed. Model and dimension are recorded, so a mismatch reports a clear error instead of silently degrading to BM25.
rag packuses embeddings. It was hardcoded to BM25 regardless of config.
Retrieval quality, measured
Built a 14-question ground-truth set and measured instead of guessing (#12):
| Config | Vector recall@10 | MRR |
|---|---|---|
nomic-embed-text @ ~2200-char chunks |
2/14 | 0.054 |
mxbai-embed-large @ ~540-char chunks |
6/14 | 0.329 |
Vector recall@200 is 10/14 — the passages are retrieved but ranked 32-141, so the remaining gap needs reranking rather than more embedding tuning. Tuning bm25_weight was measured and does not help (flat at 6/14 across 0.2-0.65).
embedding.include_path is new: prefixing embedded chunks with path:lines helps code but cost 0.12 MRR on prose.
rag ask
The agentic loop lived only in a 1,299-line example that duplicated an OpenAI client and the retriever wiring, so the CLI itself could not answer anything and new features never reached it. It is now rag ask, built on the shared adapters; the example is 172 lines. (#17)
rag ask -d ./books -q "How did Ned Stark die" --fast --hydeHyDE (--hyde) asks the LLM for a hypothetical answer and searches with that, since questions and answers do not resemble each other in embedding space. One LLM call, cached in the index, so repeats cost zero. On the test corpus it changed a wrong answer into a correct cited one at the same cost.
Every run prints what it used:
Back-and-forth rounds: 1 of 2 max
LLM calls: 1
Total tokens: 3,601 (reported by the API)
Token counts come from the provider's usage field, and the output says whether they were reported or estimated. (#14, #15)
Other fixes
rag index /tmpindexed nothing on macOS, where/tmpis a symlink.filepath.Walklstats its root, so a symlinked root yielded zero files while the CLI's own directory check passed — a silent no-op. (#13).envis read automatically, searched upward from the target and working directories. Environment variables still win. (#16)- Generation is hosted-only. No local LLM provider, in the adapter or the example; a test enforces it.
--semanticand--lexicalon bothqueryandask, and they now reject being combined. (#18)- Default config had two corrupted globs (
***.py,**/node_modulesvendor/**), so any project without arag.yamlindexed almost nothing.
Known limitations
- Recall@10 of 6/14 is honest, not good. The answers sit at rank 32-141; a reranking stage over a deep candidate pool is the next real improvement.
- The agentic loop's context-sufficiency judge is too lenient — it reports "sufficient" on context that yields a wrong answer, so
--max-iters 2often behaves like--fastplus one wasted call.
Note: this tag is v0.2.0 rather than v0.2, so it is valid semver for Go modules.
v0.1 - Initial Release
RAG CLI - Initial Release
A local-first RAG (Retrieval-Augmented Generation) toolkit for indexing and searching codebases.
Features
- BM25 + MMR retrieval pipeline - Full-text search with diversity reranking
- Two-stage hybrid search - BM25 candidates with vector reranking (#6)
- WebAssembly support - In-browser search capabilities (#8)
- Clean hexagonal architecture - Usecases depend on ports only (#7)
- Auto-indexing - With
--no-auto-indexflag to disable (#5)
CLI Commands
rag index [path] # Index files
rag query -q "search" # Search with BM25 + MMR
rag pack -q "question" -b 4000 # Pack context for LLM consumptionConfiguration
Supports YAML configuration (rag.yaml) for BM25 parameters, embedding providers (Ollama/OpenAI), and token budgets.