Releases: Gjusev/clef-evals
Releases · Gjusev/clef-evals
Release list
clef-evals v0.2.0: production quality
Production-grade rewrite of the calibration-first Clef evaluation toolkit.
Core
- Typed public API: ClefJudge / AsyncClefJudge over the real Clef schema (noul, choice, score questions with per-option probabilities)
- ClefClient: retries with exponential backoff + jitter, Retry-After support, configurable timeouts, structured ClefError hierarchy, sync + async (httpx)
- Config from environment with one-shot validation (reports all problems at once)
- Metrics: accuracy, ECE, binary + multiclass Brier, latency p50/p95/p99, cost per 1k calls from token usage
- Real CLI (clef-eval run) with JSON output and gate exit codes
Quality
- 122 unit tests, >90% coverage enforced, integration tests excluded by default
- CI matrix 3.10/3.11/3.12, ruff clean, py.typed shipped
- Regression-gate reusable GitHub Action (current vs committed baseline)
- evals/ benchmark harness + committed datasets + published reference numbers
- Kaggle kernel for cloud reproduction, calibration walkthrough notebook, showcase video renderer
clef-evals v0.1.0
Calibration-first evaluation toolkit for Cloudflare's Clef decision models.
- ClefJudge: run Clef as LLM-as-judge over eval sets
- ece() / brier_score(): calibration metrics
- clef-eval CLI entry point