Skip to content

ContextBench v1.0.0

Latest

Choose a tag to compare

@mateoosoriodelhonte mateoosoriodelhonte released this 25 Aug 22:51
36b891f

ContextBench v1.0.0 is the first complete release of the local-first RAG retrieval evaluation workbench.

Included

  • Ingest TXT, Markdown, PDF, and common source files while keeping page, heading, character-range, and chunk provenance.
  • Build frozen Qdrant local indexes with fixed-token, paragraph-aware, or heading-aware chunks.
  • Inspect vector, BM25, reciprocal rank fusion, and optional cross-encoder rankings without mixing native score scales.
  • Create versioned relevance judgments and measure Recall@K, Precision@K, MRR, Hit Rate@K, and nDCG@K.
  • Store, export, and compare immutable experiment results.
  • Use the SolidJS workbench or CLI through the same Python retrieval and evaluation code.
  • Run a zero-download deterministic demo, use consent-gated local embedding and reranker models, or ask an existing loopback Ollama model for cited answers.

Verification

The release SHA passed the Python 3.12, SolidJS, and complete Playwright browser-flow jobs in GitHub Actions. Local release checks passed 31 Python tests and 3 frontend tests, strict mypy, Ruff, Prettier, ESLint, TypeScript, package builds, the extracted-wheel browser check, and production dependency audits. The built wheel contains the browser app.

Tested limits

The deterministic hash provider, Qdrant local, BM25, hybrid retrieval, evaluation, package, and browser paths were run. Real BGE weights, cross-encoder weights, and a live Ollama model were not downloaded or run. ContextBench never downloads an Ollama model. This GitHub release contains the source archive; it is not a PyPI publication.

Run it

git clone https://github.com/mateoosoriodelhonte/contextbench.git
cd contextbench
uv sync --extra dev
cd frontend
npm ci --ignore-scripts
npm run build
cd ..
uv run contextbench serve

Open http://127.0.0.1:8000/projects and choose Load real demo.