Skip to content

Releases: Gjusev/clef-evals

clef-evals v0.2.0: production quality

Choose a tag to compare

@Gjusev Gjusev released this 01 Oct 20:45

Production-grade rewrite of the calibration-first Clef evaluation toolkit.

Core

  • Typed public API: ClefJudge / AsyncClefJudge over the real Clef schema (noul, choice, score questions with per-option probabilities)
  • ClefClient: retries with exponential backoff + jitter, Retry-After support, configurable timeouts, structured ClefError hierarchy, sync + async (httpx)
  • Config from environment with one-shot validation (reports all problems at once)
  • Metrics: accuracy, ECE, binary + multiclass Brier, latency p50/p95/p99, cost per 1k calls from token usage
  • Real CLI (clef-eval run) with JSON output and gate exit codes

Quality

  • 122 unit tests, >90% coverage enforced, integration tests excluded by default
  • CI matrix 3.10/3.11/3.12, ruff clean, py.typed shipped
  • Regression-gate reusable GitHub Action (current vs committed baseline)
  • evals/ benchmark harness + committed datasets + published reference numbers
  • Kaggle kernel for cloud reproduction, calibration walkthrough notebook, showcase video renderer

clef-evals v0.1.0

Choose a tag to compare

@Gjusev Gjusev released this 01 Oct 20:18

Calibration-first evaluation toolkit for Cloudflare's Clef decision models.

  • ClefJudge: run Clef as LLM-as-judge over eval sets
  • ece() / brier_score(): calibration metrics
  • clef-eval CLI entry point