Skip to content

v0.3.0

Latest

Choose a tag to compare

@github-actions github-actions released this 30 Sep 17:34
· 2 commits to main since this release
84c1cb6

What's Changed

  • fix: preserve falsy generation args in HELM adapter by @mrshu in #102
  • fix: correct save_to_file creating directory at file path by @mrshu in #100
  • Allow setting HELM leaderboard version by @yifanmai in #106
  • Add ARC-AGI, BFCL, SciArena, and AA (LLM only) adapters by @tommasocerruti in #98
  • feat: add base script for displaying statistics and creating visuals by @gbemike in #109
  • feat: add AlpacaEval leaderboard converter by @karthikchundi-commits in #107
  • Adapter for SWE-Bench Verified leaderboard evaluation results by @jatinganhotra in #104
  • Add CocoaBench aggregate adapter by @tommasocerruti in #111
  • Add multi swe bench and polybench leaderboards by @jatinganhotra in #110
  • Add AIR-Bench leaderboard to HELM adapter by @yifanmai in #108
  • Be robust to missing model deployment in helm results by @Erotemic in #113
  • Add LLM Stats adapter by @tommasocerruti in #118
  • Add HELM Safety leaderboard to HELM adapter by @yifanmai in #120
  • Helm stress test fixes by @Erotemic in #117
  • Tests for Inspect adapter edge cases (empty choices, local paths, misleading dataset names) by @MattFisher in #105
  • Adding support for HAL by @AsafYehudai in #114
  • feat: add Vals.ai adapter by @mrshu in #124
  • Add OpenEval adapter by @mrshu in #122
  • Update HELM adapter to use schema v0.2.2 by @yifanmai in #127
  • [ACL Shared Task] Add MMLU Pro adapter by @ajmeek in #130
  • [ACL Shared Task] Add MT-Bench adapter by @ajmeek in #128
  • [ACL Shared Task] Add HLE adapter by @ajmeek in #129
  • feat: add visualisation function by @gbemike in #121
  • Improve actions tests by @Erotemic in #135
  • Add documentation for infra by @gbemike in #146
  • docs: clarify HF upload PR workflow by @eliyahabba in #149
  • fix docs pages baseurl and readme link by @tommasocerruti in #150
  • Release workflow by @Erotemic in #151
  • Prepare prerelease 0.2.3rc1 by @Erotemic in #153
  • ci: pass repository when creating GitHub releases by @Erotemic in #155
  • adding hf community evals converter by @nelaturuharsha in #185
  • Recolor docs accent to EvalEval blue by @evijit in #186
  • fix(lm_eval): handle 'N/A' stderr from non-bootstrapped metrics by @borgr in #193
  • Fix HELM EEE instance metric rows by @Erotemic in #132
  • feat: add Mercor-Eval adapter by @tommasocerruti in #199
  • Fix LLM Stats evaluator provenance by @tommasocerruti in #136
  • Cite the paper, not the schema release by @borgr in #211
  • Cap nltk<3.10.1 in helm extra (fixes 'flaky' full/loose CI) by @borgr in #213
  • Update to validator by @nelaturuharsha in #212
  • Remove Jekyll documentation site by @nelaturuharsha in #215
  • Fix helper schema loading in installed package by @nelaturuharsha in #216
  • docs: note nltk import-guard interaction with the helm extra by @borgr in #217
  • chore: regenerate Pydantic types by @github-actions[bot] in #214
  • Refactor adapters, maintainer tools, and validator package by @nelaturuharsha in #218
  • chore: regenerate Pydantic types by @github-actions[bot] in #219
  • Add eee-dataset-conversion agent skill + a CI test that keeps it current by @borgr in #208
  • Ignore .claude/settings.local.json by @borgr in #231
  • Test that the all extra installs every optional extra by @borgr in #224
  • AGENTS.md: tie-breaker principles, and keep change-rationale out of docstrings by @borgr in #232
  • AGENTS.md: three tie-breaker principles by @borgr in #233
  • Add LEXam public leaderboard converter by @JoelNiklaus in #160
  • Add Vectara Hallucination Leaderboard adapter by @mohammadrezakarami in #157
  • Split contributor docs into CONTRIBUTING.md, and write down the review policy by @borgr in #234
  • Show validator warnings on files that pass by @borgr in #221
  • Revert "Show validator warnings on files that pass" by @nelaturuharsha in #239
  • [Feature] Add standalone local folder validator by @nelaturuharsha in #238
  • [Feature] Add EEE datastore PR review skill by @nelaturuharsha in #240
  • Rename HELM test fixture dirs to remove Windows-incompatible colons by @nelaturuharsha in #248
  • [Feature] Daily adapter ingestion cron by @Solus-QE in #249
  • Request datastore validation after each full cron submission by @evijit in #252
  • Fix LLM Stats and MMLU-Pro scheduled ingestion by @tommasocerruti in #254
  • Commit passing cron output straight to the datastore by @evijit in #255
  • Add Papers with Code adapter by @borgr in #209
  • Retry a commit the Hub refused while another adapter job held the lock by @borgr in #263
  • helm: drop the codec import and relax the nltk pin, fixing the extra on Python 3.14 by @borgr in #243
  • Test that a converted record actually passes the merge gate by @borgr in #244
  • Warn when metric_id names no metric at all by @borgr in #222
  • Commit raw store state against the live head, not a pinned parent by @borgr in #265
  • Revert the metric-identity warning until converters emit metric ids by @borgr in #266
  • Warn when a record's identity and its directory address different models by @borgr in #223
  • docs: state which root each output-dir parameter expects by @borgr in #227
  • Add Open Medical-LLM Leaderboard adapter by @borgr in #204
  • Sort the open_medical_llm imports the merge left unsorted by @borgr in #267
  • Fix 412 commit conflict handling and retry strategy in raw store ingestion by @evijit with @Copilot in #264
  • Run ruff in CI, and delete the note that said nothing did by @borgr in #273
  • Cut skill text that records how a rule was reached, not what to follow by @borgr in #274
  • Write down which name the datastore's developer folder comes from by @borgr in #271
  • fix(mt_bench): keep paper citation out of source data by @reacher-z in #257
  • Add WILD-raw item-level adapter (every_eval_ever/adapters/wild) by @borgr in #203
  • Add BenchPress score-matrix adapter by @borgr in #197
  • Convert BountyBench run logs through an adapter, with per-configuration record identity by @borgr in #188
  • Stop the type-regeneration job opening a PR about its own clock by @borgr in #278
  • Give the Terminal-Bench metric a join key by @borgr in #277
  • Put a rescaled score's reported spread on the score's scale by @borgr in #276
  • Declare the scale you computed on, rather than converting toward a guess by @borgr in #280
  • Point the developer patterns at publishing namespaces, not parent companies by @borgr in #281
  • [Submission] Add tau-bench leaderboard adapter by @benshi34 in #192
  • Tag BenchPress metric direction and split its unconvertible-row failure reasons by @borgr in #279
  • Name the benchmark, the metric and the result as three separate things by @borgr in #245
  • Report metric bounds, uncertainty and ids the way each harness computed them by @borgr in #246
  • Skip a scheduled run when source, schema and adapter are all unchanged by @borgr in #261
  • Fix HELM test fixture paths left on pre-#248 colon names by @borgr in #282
  • [Feature] Add reviewed PwC DrugBank protocol adapter by @yananlong in #260
  • [Docs] Stop listing a gate item the gate no longer checks by @borgr in #288
  • [Adapter] AlpacaEval 1.0 and 2.0, consolidated with the merged adapter and resolved through the eval-card-registry by @karthikchundi-commits in #190
  • fix(schema): do not invent max attempts by @reacher-z in #258
  • Add the Flat rebuild cron by @nelaturuharsha in #291
  • Add aiXamine adapter by @fatihdeniz in #283
  • Fix recurring Flat rebuild and Adapter ingestion failures; ingest every Terminal-Bench version; add LiveBench adapter by @evijit in #300
  • Handle comma-separated adapter input in ingestion plan job by @evijit with @Copilot in #301
  • [Submission] Add Terminal-Bench-Science adapter by @StevenDillmann in #296
  • livebench: one namespace spelling per publisher; drop resolved flat conflicts by @evijit in #303
  • Prepare release 0.3.0 by @evijit in #302

New Contributors

Full Changelog: v0.2.2...v0.3.0