Skip to content

2.20.3

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 28 Aug 13:47
· 175 commits to main since this release

2.20.3 (2026-08-28)

Ci

  • ci: limit PR dataset checks to added task files (#5293)

  • ci: limit PR dataset checks to added task files

  • ci: narrow dataset fallback fix


Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (dff378f)

Fix

  • fix: Remove legacy from benchmark names (#5310)

fix: Remove legacy and add superseeded by

Since we have the "newer version" below I think it is better to remove legacy. If we keep legacy then it might be nice to add a rule to check that superseeded by is specified.

For jpn it seems to be legacy, but I can't see any newer version?

For farsi I added superseeded by. (10d8943)

Unknown

  • Add DNA-VL-STEER-2B (#5313)

  • model: Add DNA-VL-STEER-2B

DNA-VL-STEER-2B is a language-bias-calibrated variant of
Qwen/Qwen3-VL-Embedding-2B. Identical architecture and loader, so the meta
inherits the base entry's capacity fields; it differs in declaring all 36
calibrated languages.

  • Set adapted_from instead of describing the base model in a comment (293e9ad)

  • dataset: add Spanish Wikinews clustering tasks (#5308)

  • Add Spanish Wikinews clustering tasks

  • Add Spanish Wikinews descriptive statistics

  • chore: allow Spanish prompt token in typos

  • fix: format Spanish Wikinews citations


Co-authored-by: clemente Ranokau <clemente@ranokau.com> (5b35ec8)

  • model: Add amgix/static-retrieval-multilingual-69m-v1 (#5312)

Add amgix/static-retrieval-multilingual-69m-v1 to misc_models (a92ff3e)

  • Add dense webvid retrieval (#5215)

  • First commit dense_webvid_retrieval

  • Fix revision and pass test


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (775411c)

  • Create speech edit acoustic AT2A dataset (#5285)

  • Create speech edit acoustic AT2A dataset

Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>

  • update main_score

Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>

  • update bibtex

Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>

  • Update reference link to arxiv (abs)

Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>

  • update bibtex to pass test

  • some nits and license fix

  • Add SpeechEditAcousticRetrieval to duplicate_text

  • uv ruff fix

  • Update data_prep.py

  • Update test_task_quality.py


Co-authored-by: nehal309 <20415842+nehal309@users.noreply.github.com> (be7d3a1)

  • Add ACM composed audio dataset (#5187)

  • ACM dataset

  • fix test_task_quality.py

  • Update test_task_quality.py

  • Update init.py

  • Apply suggestion from @Samoed


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (8e9f956)

  • Add narcolepticchicken/octen-law-8b-v1 (#5305)

  • Add Octen Law 8B v1 model

  • Disclose model training datasets

  • style: format Octen model registration

  • Use ScoringFunction enum for Octen Law

  • Remove mock run results

  • Update Octen Law namespace to Litil Labs (9d2de3a)

  • Add NSynth instrument family clustering task (#5111)

  • task: add NSynth instrument family clustering

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

  • test: allow upstream NSynth audio duplicates in clustering task

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> (d7adced)

  • dataset: add VCDB core video and audio-video retrieval (#5269)

  • dataset: add VCDB core video retrieval

  • dataset: add VCDB core audio retrieval

  • dataset: document VCDB duplicate audio

  • dataset: mark VCDB license unspecified

  • dataset: replace VCDB audio task with audio-video (b398a41)

  • dataset: add ManiSkill i2v/v2i robot manipulation retrieval (MOEB) (#5257)

  • dataset: add ManiSkill i2v/v2i robot manipulation retrieval (MOEB)

Adds ManiSkillI2VRetrieval and ManiSkillV2IRetrieval, image<->video
retrieval over ManiSkill3 motion-planning demonstrations (8 tabletop
tasks, 150 successful episodes each, replayed via environment states and
rendered at 256x256 from two viewpoints: base sensor camera for
goal-state images, human render camera for videos). Per task, episodes
split into disjoint query (10) and corpus (140) pools; relevance is
task-level and multi-positive with the query's source episode held out.
Instance-level 1:1 designs over the near-duplicate corpus measured at
chance level for current models and were rejected. Adds the v2i task
category and Robotics domain (same additions as the LIBERO PR; merges
cleanly in either order).

Reference runs (ndcg@10): mteb/baseline-random-encoder 0.1236 (i2v) /
0.1129 (v2i); jinaai/jina-embeddings-v5-omni-nano 0.4193 (i2v) /
0.4766 (v2i).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3

  • fix: clarify ManiSkill relevance definition

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3

  • fix: deduplicate ManiSkill episodes by seed (official demos repeat seeds)

The official motion-planning demo files reuse episode seeds, so the
first-150 slice contained 25 duplicate corpus episodes (caught by
test_dataset_quality). Select the first 150 unique-seed successful
episodes instead, rebuild both datasets, refresh revisions, descriptive
stats and reference results.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3


Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (739d8ef)

  • dataset: add LIBERO i2v/v2i robot manipulation retrieval (MOEB) (#5255)

  • dataset: add LIBERO i2v/v2i robot manipulation retrieval (MOEB)

Adds LIBEROI2VRetrieval and LIBEROV2IRetrieval, image<->video retrieval
over the LIBERO robot manipulation benchmark (40 tasks, 1,693 episodes,
256x256 @ 10 fps). Queries are goal-state images (final frames) of
held-out episodes; relevance is task-level and multi-positive with the
query's source episode excluded from the corpus, so exact frame matching
cannot solve the task. Adds the v2i task category (previously empty
direction) and a Robotics task domain.

Reference runs (ndcg@10): mteb/baseline-random-encoder 0.0297 (i2v) /
0.0261 (v2i); jinaai/jina-embeddings-v5-omni-nano 0.3537 (i2v) /
0.3312 (v2i).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3

  • fix: clarify LIBERO relevance definition and correct license to cc-by-4.0

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3


Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (718d295)

  • Add Clotho moment dataset (AT2A) (#5287)

  • Add clotho-moment retrieval task

  • update version

Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>

  • update bibtex title

Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>

  • Add task description and fix tests

Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>

  • Move descriptive stage to right directory.

  • Fix formatting in Clotho moment retrieval file

  • Update test_task_quality.py

  • Update test_task_quality.py


Co-authored-by: nehal309 <20415842+nehal309@users.noreply.github.com> (bb05784)

  • dataset: add EDIR (#5247)

  • dataset: add EDIR

  • update bibtex format and mark known duplicates

  • add script for processing data

  • lint format

  • add prompt

  • update query prompt and delete document prompt (2440f68)

  • Update GreenLeaf Law Embed Tiny: add 35+ languages (#5304)

Update GreenLeaf languages: add 35+ supported languages

Model trained on multilingual legal corpus covering 35+ languages
including English, German, French, Spanish, Chinese, Japanese,
Korean, Arabic, Hindi, and others.

Co-authored-by: Surya-saka <surya@judicialmind.ai> (626de72)

  • [MOEB] Add UniME-V2-LLaVA-OneVision-8B model (#5296)

  • Add UniME V2 model

Signed-off-by: jupyterjazz <saba.sturua@jina.ai>

  • chore: remove tests

Signed-off-by: jupyterjazz <saba.sturua@jina.ai>

  • Delete mteb_mock_run_results.md

Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (9ef064e)