Skip to content

2.12.21

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 18 Apr 10:49
· 759 commits to main since this release

2.12.21 (2026-04-18)

Ci

  • ci: add workflow to auto-update leaderboard model list (#4402)

Adds a standalone script that generates the model list from scratch
and a CI workflow that pushes it to the HF leaderboard space weekly,
on model file changes, or via manual dispatch.

Closes #4316 (18e8e63)

Fix

  • fix: Add required_dependencies to model meta (#4356)

  • add required_dependencies to model meta

  • add extra group name

  • add to model to python

  • update handling dependencies

  • fix deps

  • fix test

  • remove usage of requires_package

  • remove image/audio dependencies

  • fixes after merge

  • add deprecated function

  • fix test

  • skip check for baseline

  • fix test

  • update lock

  • optionally check torchaudio in test (e2e7174)

Unknown

  • Remove video folder (#4424)

remove video folder (011bbf5)

update dataset card (43d1b21)

  • tests: Add test to ensure coverage of reference models (#4216)

  • Reference models tests

  • Reference models tests

  • Reference models tests

  • fix: address PR review comments for reference model tests

  • Use cache.load_results() instead of manually walking cache directories
  • Dynamically compute target benchmarks from all leaderboard benchmarks
    minus an exclusion list, so new benchmarks are automatically tested
  • Add text-only modality check for task-model compatibility
  • Filter retrieval-only models by task type AND text modalities
  • fix: use isinstance check for retrieval subtypes

Check isinstance(task, AbsTaskRetrieval) instead of string comparison
with task.metadata.type, so reranking and instruction retrieval tasks
are correctly included for retrieval-only models like bm25s.

  • fix: handle empty sim_scores in confidence_scores

Return zero confidence scores when sim_scores list is empty,
which can happen when BM25 returns no results for a query
in reranking tasks.

  • fix: address PR review comments for reference model tests
  • Remove RTEB variant exclusions to test all RTEB benchmarks
    (per Kenneth's feedback to include the full RTEB set)
  • fix: use benchmark_selector.py as source of truth for leaderboard benchmarks

Address Kenneth's review comments:

  • Use GP_BENCHMARK_ENTRIES + R_BENCHMARK_ENTRIES from benchmark_selector.py
    instead of display_on_leaderboard flag (which includes benchmarks not
    actually shown on the leaderboard)
  • Clean up EXCLUDED_BENCHMARKS to only contain actual leaderboard benchmarks
    (multimodal ones that text-only reference models can't run)
  • Remove RTEB variant exclusions to test the full RTEB set
  • fix: remove all benchmark exclusions, rely on task-level filtering

Task-level filtering (_is_text_only_task, RETRIEVAL_ONLY_MODELS) already
handles model-task compatibility. No need to exclude entire benchmarks —
non-text tasks within multimodal benchmarks are skipped automatically.

  • fix: use display_on_leaderboard flag now that PR #4288 is merged

Simplify _get_target_benchmarks to use display_on_leaderboard=True,
which now correctly reflects the actual leaderboard (fixed in #4288).
Remove benchmark_selector imports and exclusion list — task-level
filtering handles model-task compatibility.

  • fix: pass Benchmark objects directly instead of names

Address Samoed's review: use Benchmark objects in parametrize
instead of looking up by name twice.

  • speedup test

  • fix issue with aggregate

  • fix: address review - reuse _check_model_modalities, trim workflow triggers

  • fix: restore TARGET_BENCHMARKS definition, remove stale _get_target_benchmarks call

  • fix: inline modality check to avoid private import, filter image-only tasks

  • fix: use strict modality subset check to exclude image/multimodal tasks

  • fix: restore RETRIEVAL_ONLY_MODELS for BM25 task filtering

  • fix: add mteb/benchmarks/** to workflow triggers


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4a90b28)

  • [MVEB] Adding WorldSense1Min Task (Clustering) (#4393)

  • [MVEB] Adding WorldSense1Min Task (Clustering)

  • remove local test

  • Update mteb/tasks/init.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • removing stats

  • moving video clustering tasks to clustering

  • uncomment Video task

  • add results

  • update license

  • remove results


Co-authored-by: wissam-KH <wissam.siblini@komodohealth.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (f46cb7b)

  • [MVEB] Adding AVE-Dataset Task (Clustering) (#4416)

  • [MVEB] Adding AVE-Dataset Task (Clustering)

  • uncomment video clustering task

  • remove results (61e7f3f)

  • tests: add regression test for double loading (#4407)

add regression test (e946e1e)

  • add HMDB51 dataset (#4398)

  • add HMDB51 dataset

  • update

  • Update mteb/abstasks/task_metadata.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Update mteb/tasks/classification/eng/hmdb51_classification.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • fix lint

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (22bc680)