Releases: vishwanathakuthota/openvals
Releases · vishwanathakuthota/openvals
Release list
0.5.5
What's Changed
- Fix dataset validation to support input and output keys by @kreshnaaa in #4
New Contributors
- @kreshnaaa made their first contribution in #4
Full Changelog: v0.5.0...v0.5.5
v0.5.0
v0.4.0
161201 experiments with Fact checking
Full Changelog: v0.2.0...v0.3.0
v0.2.0
v0.1.5
v0.1.0
🚀 OpenVals
Open-source AI Model Evaluation, Benchmarking & Trust Layer for LLMs
OpenVals helps you evaluate, compare, and trust AI models (OpenAI, Claude, Gemini, Ollama, open-source models) using standardized metrics + real-world benchmarking + explainable scoring.
⸻
Why OpenVals?
Most AI models look great in demos—but fail in production.
OpenVals answers:
- Which model is actually best for your use case?
- How do models compare beyond just “accuracy”?
- Can I trust this model in production?
- Which model is fastest, safest, and most reliable?
🧠 What OpenVals Does
OpenVals is a 3-layer AI evaluation system:
- 📊 Evaluation Layer
- Accuracy scoring
- Semantic similarity
- Latency measurement
- Reliability scoring
- Safety scoring
- 🏁 Benchmarking Layer
- Multi-model comparison (Ollama, OpenAI, Claude, etc.)
- Normalized scoring across models
- Ranking engine
- 🎯 Recommendation Engine
- Suggests best model for your dataset
- Tradeoff-aware ranking (speed vs accuracy vs safety)
- Use-case based model selection (coming next version)
Full Changelog: v0.0.5...v0.1.0
