Skip to content

2.12.22

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 18 Apr 23:08
· 753 commits to main since this release

2.12.22 (2026-04-18)

Fix

  • fix: KeyError on aggregated tasks with eval_langs (#4439)

Fix KeyError on aggregated tasks with dict eval_langs

When aggregated tasks (e.g. VisualSTS17Multilingual) have eval_langs
as a dict, hf_subsets_to_langscripts lacks a "default" key. The
aggregated score uses "default" as subset, causing a KeyError in
TaskResult.from_task_results. Fall back to collecting all languages
from the mapping when the subset key is missing.

Closes #4437

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> (5fc2867)

Unknown

  • Remove skip_first for jamaltartistit (#4435) (16ba72a)

  • Fix: apply skip_first_result when computing hit_rate metric (#4427)

  • Fix skip_first_result not applied to hit_rate metric

  • lint


Co-authored-by: Rakshitha Ireddi <rakshithaireddi@Rakshithas-MacBook-Pro.local>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (8570b74)

  • [MVEB] Adding MUSIC-AVQA Task (Clustering) (#4426)

  • [MVEB] Adding MUSIC-AVQA Task (Clustering)

  • simplify description (04f2f4b)

  • [MVEB] Add Breakfast video classification task (#4431)

  • [MVEB] Add Breakfast video classification task\n\nAdd BreakfastClassification task for the Breakfast Actions dataset (Kuehne et al., CVPR 2014). The dataset contains 433 videos of 10 breakfast-related activities recorded in 18 kitchens. Uses 5-fold cross-validation since the dataset only has a test split.\n\nRandom baseline accuracy: 0.1247 (near-random for 10 classes).\n\nAddresses part of #4130 (MVEB Overview - Classification).

  • lint


Co-authored-by: Yashwanth Devavarapu <yashwanthdevavarapu@Yashwanths-MacBook-Pro.local>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (98d682c)

  • Add ActivityNet_Captions_val2 video retrieval tasks (V2T and T2V) (#4429)

  • Add ActivityNet_Captions_val2 video retrieval tasks (V2T and T2V)

Made-with: Cursor

  • Fix all sort order for isort-style linting (RUF033)

Made-with: Cursor

  • Update mteb/tasks/retrieval/eng/activitynet_captions_t2v_retrieval.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Consolidate ActivityNet Captions retrieval tasks into a single file

Made-with: Cursor


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (7539339)

  • add mteb/DiDeMo dataset (#4425)

  • add mteb/MSVD dataset

  • add mteb/DiDeMo dataset

  • add comb

  • update

  • Consolidate DiDeMo retrieval tasks into a single file

Merge 4 separate DiDeMo task files into didemo_retrieval.py with a
shared _load_didemo helper, reducing duplication while preserving
all task names and metadata.

Made-with: Cursor

  • Remove unused Dataset import from didemo_retrieval

Made-with: Cursor (2e00e5b)

  • add mteb/TUNA-Bench_1K dataset (#4428)

Add TUNA-Bench_1K video retrieval tasks (V2T and T2V)

Made-with: Cursor (b14dcc5)

  • Update vllm_wrapper.py (#4418)

fix compatibility with newer vllm versions (5ef64c3)