Drop-in LLMOps observability. Self-hosted. SQLite-default. Persona-aware. Multi-judge consensus.
Version: v0.1.10 | pip install omnismart-rageval
🔗 Live dashboard: https://rageval.ysiddo-ai-projects.app/ · browser dashboard (score a query + view metrics). Also fully scriptable — API:
/health,/eval/*viacurl/HTTPie. On-demand backend (first request ~30–60 s to wake). Self-hosting: see SELF_HOSTING.md.
from rageval import track
@track(model="anthropic/claude-sonnet-4-6", persona="cfo")
async def answer_question(query: str, context_chunks: list[str]) -> str:
...That's it. Open the dashboard at localhost:8003.
| Metric | Definition |
|---|---|
| Retrieval relevance | Cosine sim between query and retrieved chunks (BGE-large by default) |
| Groundedness consensus | Multi-judge LLM scoring (Claude Haiku 4.5 + Groq LLaMA 3.3 + GPT-5-mini), flags disagreement |
| Faithfulness | Per-sentence max-similarity to any chunk (NLI proxy) |
| Cost | USD per interaction, tracked by model |
| Latency | End-to-end wall-clock |
| Feature | RAGeval | Phoenix | Langfuse | TruLens |
|---|---|---|---|---|
| Self-hosted | ✅ | ✅ | ✅ | ✅ |
| SQLite default | ✅ | ❌ | ❌ | ❌ |
| Drop-in decorator | ✅ | partial | ❌ | partial |
| Persona-aware | ✅ | ❌ | ❌ | ❌ |
| Multi-judge consensus | ✅ | ❌ | ❌ | ❌ |
| Cost tracking | ✅ | ✅ | ✅ | partial |
| Setup time | 60 sec | 10 min | 15 min | 10 min |
pip install --index-url https://gateway.ysiddo-ai-projects.app/pypi/simple/ omnismart-rageval # distribution name; CLI + import remain `rageval`
rageval init # creates ~/.rageval/rageval.db
rageval serve --port 8003from rageval import track
@app.post("/ask")
@track(model="anthropic/claude-sonnet-4-6", persona="cfo")
async def ask(query: str):
chunks = await retriever.search(query)
return await llm.generate(query, chunks=chunks)@track(model="groq/llama-3.3-70b-versatile")
def chain_invoke(query: str, context_chunks: list[str]):
return chain.invoke({"query": query, "context": context_chunks})| Method | Path | Purpose |
|---|---|---|
| GET | /health | Liveness |
| POST | /eval/log | Score + store |
| POST | /eval/score | Score only (no storage) |
| GET | /eval/metrics?days=7 | Aggregate dashboard data |
| GET | /eval/queries | Query log (filter by needs_review) |
| GET | /eval/cost-report?days=30 | Cost breakdown by day + model |
| GET | /eval/alerts | Recent flagged queries |
| POST | /eval/retrieval-bench | A/B compare retrieval strategies |
| POST | /eval/embedding-comparison | Compare embedding models |
38 test functions across smoke, API, evaluator, decorator, DSPy integration, store, and e2e:
pytest tests/ -qAGPL-3.0
This project is open-source under the AGPL-3.0 License. It is completely free for researchers, students, and open-source hobbyists.
Commercial Use: The AGPLv3 license requires that any proprietary network service (SaaS, internal corporate tools) that uses or modifies this code must also open-source its entire backend.
If you wish to use this framework in a closed-source commercial environment, or require Enterprise features (SSO, Active Directory, Custom VPC Deployment, Strict RBAC), you must obtain a Commercial License. Please reach out to discuss commercial licensing and integration consulting.
This project collects anonymous, GDPR-compliant startup pings to help the author understand usage volume and prioritize development.
- What is collected: A startup event timestamp and anonymized deployment origin. No API keys, no user prompts, and no sensitive application data is ever collected.
- How to disable: We respect your privacy and development environment. To opt-out, simply set
TELEMETRY_OPT_OUT=truein your.envfile.
This project is licensed under the AGPL-3.0 License.
Commercial Use: If you wish to use this software commercially without releasing your own source code, please see COMMERCIAL.md to obtain a commercial license.
Telemetry: See TELEMETRY.md for our privacy-first data practices.