Skip to content

Releases: aifabrice/jev-rag

Jev RAG v0.3.0 — Agentic and hybrid retrieval

Choose a tag to compare

@aifabrice aifabrice released this 26 Sep 15:28

What’s new

Jev RAG 0.3.0 expands the local-first retrieval stack while keeping BM25 + Jev as the vector-free default.

  • Agentic lexical retrieval: two-round MiniMax query planning, multi-query reciprocal-rank fusion, local plan caching, and defensive planner-output validation.
  • Optional hybrid retrieval: OpenRouter embeddings, a local NumPy vector cache, and BM25 + embedding RRF before Jev reranking.
  • Zero-configuration folder discovery: jev-rag serve indexes ~/Documents by default and avoids indexing the source checkout.
  • End-to-end observability: BM25, Agentic planning, embeddings/RRF, Jev, first-token, generation, and total latency reporting.
  • Public evaluation: complete BEIR NFCorpus results plus an interactive benchmark explorer.

Recorded NFCorpus results

Complete test split: 3,633 documents and 323 queries.

Pipeline nDCG@10 MRR@10 Recall@10
BM25 top 30 0.305654 0.512697 0.147309
BM25 top 30 + Jev 0.353235 0.585817 0.158667
Agentic lexical + Jev 0.430969 0.644041 0.204138
Hybrid top 50 + Jev 0.444327 0.654583 0.214907

These are project-run results, not an official MTEB submission.

Quick start

git clone https://github.com/aifabrice/jev-rag.git
cd jev-rag
python -m pip install -e '.[documents]'
jev-rag serve

Open http://127.0.0.1:8765. Provider-backed reranking, query planning, embeddings, and answer generation require the corresponding API keys. Offline BM25 indexing and search remain local.

Jev RAG v0.2.0 — first public alpha

Choose a tag to compare

@aifabrice aifabrice released this 26 Sep 05:36

Jev RAG v0.2.0 is the first public alpha release of a vector-free local knowledge retrieval stack. It combines SQLite FTS5/BM25 retrieval, Jev evidence reranking, and optional grounded streaming answers through OpenRouter.

Highlights

  • Index a local folder without embeddings, a vector database, or a GPU.
  • Retrieve up to 30 BM25 candidates and rerank them with batched Jev requests.
  • Stream cited MiniMax answers while measuring retrieval, reranking, first-token, generation, and total latency.
  • Choose automatic, paragraph, or no-chunking modes.
  • Load Markdown, text, HTML, JSON, CSV, YAML, DOCX, and PDF documents.
  • Run a reproducible BM25/Jev retrieval benchmark with your own JSONL cases.
  • Use the bilingual English and Simplified Chinese documentation.

Try it

python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[documents]"
cp .env.example .env
python local_kb.py index
python local_kb.py serve

Then open http://127.0.0.1:8765.

Important limitations

This is alpha software. Lexical retrieval may miss synonyms and paraphrases. Jev and the answer model are remote services, so queries and selected passages leave the machine when those features are enabled. The local web server has no authentication and should not be exposed directly to the public internet.

See the README, evaluation guide, and security policy before use.