What's Changed
- fix: preserve falsy generation args in HELM adapter by @mrshu in #102
- fix: correct save_to_file creating directory at file path by @mrshu in #100
- Allow setting HELM leaderboard version by @yifanmai in #106
- Add ARC-AGI, BFCL, SciArena, and AA (LLM only) adapters by @tommasocerruti in #98
- feat: add base script for displaying statistics and creating visuals by @gbemike in #109
- feat: add AlpacaEval leaderboard converter by @karthikchundi-commits in #107
- Adapter for SWE-Bench Verified leaderboard evaluation results by @jatinganhotra in #104
- Add CocoaBench aggregate adapter by @tommasocerruti in #111
- Add multi swe bench and polybench leaderboards by @jatinganhotra in #110
- Add AIR-Bench leaderboard to HELM adapter by @yifanmai in #108
- Be robust to missing model deployment in helm results by @Erotemic in #113
- Add LLM Stats adapter by @tommasocerruti in #118
- Add HELM Safety leaderboard to HELM adapter by @yifanmai in #120
- Helm stress test fixes by @Erotemic in #117
- Tests for Inspect adapter edge cases (empty choices, local paths, misleading dataset names) by @MattFisher in #105
- Adding support for HAL by @AsafYehudai in #114
- feat: add Vals.ai adapter by @mrshu in #124
- Add OpenEval adapter by @mrshu in #122
- Update HELM adapter to use schema v0.2.2 by @yifanmai in #127
- [ACL Shared Task] Add MMLU Pro adapter by @ajmeek in #130
- [ACL Shared Task] Add MT-Bench adapter by @ajmeek in #128
- [ACL Shared Task] Add HLE adapter by @ajmeek in #129
- feat: add visualisation function by @gbemike in #121
- Improve actions tests by @Erotemic in #135
- Add documentation for infra by @gbemike in #146
- docs: clarify HF upload PR workflow by @eliyahabba in #149
- fix docs pages baseurl and readme link by @tommasocerruti in #150
- Release workflow by @Erotemic in #151
- Prepare prerelease 0.2.3rc1 by @Erotemic in #153
- ci: pass repository when creating GitHub releases by @Erotemic in #155
- adding hf community evals converter by @nelaturuharsha in #185
- Recolor docs accent to EvalEval blue by @evijit in #186
- fix(lm_eval): handle 'N/A' stderr from non-bootstrapped metrics by @borgr in #193
- Fix HELM EEE instance metric rows by @Erotemic in #132
- feat: add Mercor-Eval adapter by @tommasocerruti in #199
- Fix LLM Stats evaluator provenance by @tommasocerruti in #136
- Cite the paper, not the schema release by @borgr in #211
- Cap nltk<3.10.1 in helm extra (fixes 'flaky' full/loose CI) by @borgr in #213
- Update to validator by @nelaturuharsha in #212
- Remove Jekyll documentation site by @nelaturuharsha in #215
- Fix helper schema loading in installed package by @nelaturuharsha in #216
- docs: note nltk import-guard interaction with the helm extra by @borgr in #217
- chore: regenerate Pydantic types by @github-actions[bot] in #214
- Refactor adapters, maintainer tools, and validator package by @nelaturuharsha in #218
- chore: regenerate Pydantic types by @github-actions[bot] in #219
- Add eee-dataset-conversion agent skill + a CI test that keeps it current by @borgr in #208
- Ignore .claude/settings.local.json by @borgr in #231
- Test that the
allextra installs every optional extra by @borgr in #224 - AGENTS.md: tie-breaker principles, and keep change-rationale out of docstrings by @borgr in #232
- AGENTS.md: three tie-breaker principles by @borgr in #233
- Add LEXam public leaderboard converter by @JoelNiklaus in #160
- Add Vectara Hallucination Leaderboard adapter by @mohammadrezakarami in #157
- Split contributor docs into CONTRIBUTING.md, and write down the review policy by @borgr in #234
- Show validator warnings on files that pass by @borgr in #221
- Revert "Show validator warnings on files that pass" by @nelaturuharsha in #239
- [Feature] Add standalone local folder validator by @nelaturuharsha in #238
- [Feature] Add EEE datastore PR review skill by @nelaturuharsha in #240
- Rename HELM test fixture dirs to remove Windows-incompatible colons by @nelaturuharsha in #248
- [Feature] Daily adapter ingestion cron by @Solus-QE in #249
- Request datastore validation after each full cron submission by @evijit in #252
- Fix LLM Stats and MMLU-Pro scheduled ingestion by @tommasocerruti in #254
- Commit passing cron output straight to the datastore by @evijit in #255
- Add Papers with Code adapter by @borgr in #209
- Retry a commit the Hub refused while another adapter job held the lock by @borgr in #263
- helm: drop the codec import and relax the nltk pin, fixing the extra on Python 3.14 by @borgr in #243
- Test that a converted record actually passes the merge gate by @borgr in #244
- Warn when metric_id names no metric at all by @borgr in #222
- Commit raw store state against the live head, not a pinned parent by @borgr in #265
- Revert the metric-identity warning until converters emit metric ids by @borgr in #266
- Warn when a record's identity and its directory address different models by @borgr in #223
- docs: state which root each output-dir parameter expects by @borgr in #227
- Add Open Medical-LLM Leaderboard adapter by @borgr in #204
- Sort the open_medical_llm imports the merge left unsorted by @borgr in #267
- Fix 412 commit conflict handling and retry strategy in raw store ingestion by @evijit with @Copilot in #264
- Run ruff in CI, and delete the note that said nothing did by @borgr in #273
- Cut skill text that records how a rule was reached, not what to follow by @borgr in #274
- Write down which name the datastore's developer folder comes from by @borgr in #271
- fix(mt_bench): keep paper citation out of source data by @reacher-z in #257
- Add WILD-raw item-level adapter (every_eval_ever/adapters/wild) by @borgr in #203
- Add BenchPress score-matrix adapter by @borgr in #197
- Convert BountyBench run logs through an adapter, with per-configuration record identity by @borgr in #188
- Stop the type-regeneration job opening a PR about its own clock by @borgr in #278
- Give the Terminal-Bench metric a join key by @borgr in #277
- Put a rescaled score's reported spread on the score's scale by @borgr in #276
- Declare the scale you computed on, rather than converting toward a guess by @borgr in #280
- Point the developer patterns at publishing namespaces, not parent companies by @borgr in #281
- [Submission] Add tau-bench leaderboard adapter by @benshi34 in #192
- Tag BenchPress metric direction and split its unconvertible-row failure reasons by @borgr in #279
- Name the benchmark, the metric and the result as three separate things by @borgr in #245
- Report metric bounds, uncertainty and ids the way each harness computed them by @borgr in #246
- Skip a scheduled run when source, schema and adapter are all unchanged by @borgr in #261
- Fix HELM test fixture paths left on pre-#248 colon names by @borgr in #282
- [Feature] Add reviewed PwC DrugBank protocol adapter by @yananlong in #260
- [Docs] Stop listing a gate item the gate no longer checks by @borgr in #288
- [Adapter] AlpacaEval 1.0 and 2.0, consolidated with the merged adapter and resolved through the eval-card-registry by @karthikchundi-commits in #190
- fix(schema): do not invent max attempts by @reacher-z in #258
- Add the Flat rebuild cron by @nelaturuharsha in #291
- Add aiXamine adapter by @fatihdeniz in #283
- Fix recurring Flat rebuild and Adapter ingestion failures; ingest every Terminal-Bench version; add LiveBench adapter by @evijit in #300
- Handle comma-separated adapter input in ingestion plan job by @evijit with @Copilot in #301
- [Submission] Add Terminal-Bench-Science adapter by @StevenDillmann in #296
- livebench: one namespace spelling per publisher; drop resolved flat conflicts by @evijit in #303
- Prepare release 0.3.0 by @evijit in #302
New Contributors
- @tommasocerruti made their first contribution in #98
- @gbemike made their first contribution in #109
- @karthikchundi-commits made their first contribution in #107
- @jatinganhotra made their first contribution in #104
- @MattFisher made their first contribution in #105
- @AsafYehudai made their first contribution in #114
- @ajmeek made their first contribution in #130
- @eliyahabba made their first contribution in #149
- @borgr made their first contribution in #193
- @JoelNiklaus made their first contribution in #160
- @mohammadrezakarami made their first contribution in #157
- @Solus-QE made their first contribution in #249
- @evijit with @Copilot made their first contribution in #264
- @reacher-z made their first contribution in #257
- @benshi34 made their first contribution in #192
Full Changelog: v0.2.2...v0.3.0