Releases: JarJarBeatyourattitude/evalt
Release list
Evalt 0.9.5
Evalt 0.9.5 makes the tested OpenRouter Chat Completions request envelope part of each durable route. Structured output, tools, sampling controls, provider routing, plugins, multimodal message content, and future JSON request fields now run under the same settings during evaluation and production. Overrides emit a typed warning and lose the saved quality claim; strict mode stops them before provider spend. Tool-call-only responses are preserved as structured results.
Evalt 0.9.1
Automatic first-route tournaments now bound stalled providers to a 120-second default, prune dominated higher reasoning once a lower effort passes, and print broad-screen/model progress with validation, latency, spend, and elapsed time. The full live acceptance designed 25 cases, calibrated the judge, tested five models and 21 prompt packages, promoted a 100% final-test route, and reused it on the next call.
Evalt 0.9.0
Evalt now builds the first production route automatically: AI-designed test cases, calibrated judging, model/reasoning/prompt/few-shot search, an untouched final-test gate, durable route reuse, and explicit evidence provenance. The first real input is design context without an approved label. Legacy bootstrap-only behavior remains available with first_run=" bootstrap.
Evalt 0.8.22
Makes bootstrap status unambiguous, shows feedback accumulation, launches the first bounded tournament after feedback is recorded, preserves in-flight maintenance for short scripts, isolates evidence after prompt changes, and adds budget-shared AI suite design with review and explicitly labeled autopilot modes.
Evalt 0.8.21
Fixes the hidden .02 production-call ceiling when price_usd is omitted, keeps test_budget_usd separate, and adds compact interactive stderr progress plus structured progress callbacks. Includes a regression for the reported .057045 request.