Decision Benchmarks
0.4.0 - 2026-09-25
The decision benchmark.
Added
- The decision benchmark 1.0 (
thinkless bench decisions): eight public
tasks across the four question kinds (banking77, clinc150 with out of
scope, MASSIVE in five languages, MultiWOZ 2.2 conversations, jailbreaks,
Civil Comments toxicity, HelpSteer2 helpfulness ratings and WNUT 2017
entities), with frozen rows and published hashes, calibration rows from a
separate split, result files that hold every prediction,verifythat
recomputes every metric, a generated leaderboard and open submissions.
Baselines for GLiNER 2.5, Laya, Qwen3-1.7B and Qwen 3.7 Flash.
Changed
- A list of chat messages now renders as
role: textlines for providers
that read text, instead of onekey: valueline per field. - README links to the documentation point at the documentation site.