2.20.3
2.20.3 (2026-08-28)
Ci
-
ci: limit PR dataset checks to added task files (#5293)
-
ci: limit PR dataset checks to added task files
-
ci: narrow dataset fallback fix
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (dff378f)
Fix
- fix: Remove legacy from benchmark names (#5310)
fix: Remove legacy and add superseeded by
Since we have the "newer version" below I think it is better to remove legacy. If we keep legacy then it might be nice to add a rule to check that superseeded by is specified.
For jpn it seems to be legacy, but I can't see any newer version?
For farsi I added superseeded by. (10d8943)
Unknown
-
Add DNA-VL-STEER-2B (#5313)
-
model: Add DNA-VL-STEER-2B
DNA-VL-STEER-2B is a language-bias-calibrated variant of
Qwen/Qwen3-VL-Embedding-2B. Identical architecture and loader, so the meta
inherits the base entry's capacity fields; it differs in declaring all 36
calibrated languages.
-
Set adapted_from instead of describing the base model in a comment (
293e9ad) -
dataset: add Spanish Wikinews clustering tasks (#5308)
-
Add Spanish Wikinews clustering tasks
-
Add Spanish Wikinews descriptive statistics
-
chore: allow Spanish prompt token in typos
-
fix: format Spanish Wikinews citations
Co-authored-by: clemente Ranokau <clemente@ranokau.com> (5b35ec8)
- model: Add amgix/static-retrieval-multilingual-69m-v1 (#5312)
Add amgix/static-retrieval-multilingual-69m-v1 to misc_models (a92ff3e)
-
Add dense webvid retrieval (#5215)
-
First commit dense_webvid_retrieval
-
Fix revision and pass test
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (775411c)
-
Create speech edit acoustic AT2A dataset (#5285)
-
Create speech edit acoustic AT2A dataset
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
- update main_score
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
- update bibtex
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
- Update reference link to arxiv (abs)
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
-
update bibtex to pass test
-
some nits and license fix
-
Add SpeechEditAcousticRetrieval to duplicate_text
-
uv ruff fix
-
Update data_prep.py
-
Update test_task_quality.py
Co-authored-by: nehal309 <20415842+nehal309@users.noreply.github.com> (be7d3a1)
-
Add ACM composed audio dataset (#5187)
-
ACM dataset
-
fix test_task_quality.py
-
Update test_task_quality.py
-
Update init.py
-
Apply suggestion from @Samoed
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (8e9f956)
-
Add narcolepticchicken/octen-law-8b-v1 (#5305)
-
Add Octen Law 8B v1 model
-
Disclose model training datasets
-
style: format Octen model registration
-
Use ScoringFunction enum for Octen Law
-
Remove mock run results
-
Update Octen Law namespace to Litil Labs (
9d2de3a) -
Add NSynth instrument family clustering task (#5111)
-
task: add NSynth instrument family clustering
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- test: allow upstream NSynth audio duplicates in clustering task
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> (d7adced)
-
dataset: add VCDB core video and audio-video retrieval (#5269)
-
dataset: add VCDB core video retrieval
-
dataset: add VCDB core audio retrieval
-
dataset: document VCDB duplicate audio
-
dataset: mark VCDB license unspecified
-
dataset: replace VCDB audio task with audio-video (
b398a41) -
dataset: add ManiSkill i2v/v2i robot manipulation retrieval (MOEB) (#5257)
-
dataset: add ManiSkill i2v/v2i robot manipulation retrieval (MOEB)
Adds ManiSkillI2VRetrieval and ManiSkillV2IRetrieval, image<->video
retrieval over ManiSkill3 motion-planning demonstrations (8 tabletop
tasks, 150 successful episodes each, replayed via environment states and
rendered at 256x256 from two viewpoints: base sensor camera for
goal-state images, human render camera for videos). Per task, episodes
split into disjoint query (10) and corpus (140) pools; relevance is
task-level and multi-positive with the query's source episode held out.
Instance-level 1:1 designs over the near-duplicate corpus measured at
chance level for current models and were rejected. Adds the v2i task
category and Robotics domain (same additions as the LIBERO PR; merges
cleanly in either order).
Reference runs (ndcg@10): mteb/baseline-random-encoder 0.1236 (i2v) /
0.1129 (v2i); jinaai/jina-embeddings-v5-omni-nano 0.4193 (i2v) /
0.4766 (v2i).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
- fix: clarify ManiSkill relevance definition
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
- fix: deduplicate ManiSkill episodes by seed (official demos repeat seeds)
The official motion-planning demo files reuse episode seeds, so the
first-150 slice contained 25 duplicate corpus episodes (caught by
test_dataset_quality). Select the first 150 unique-seed successful
episodes instead, rebuild both datasets, refresh revisions, descriptive
stats and reference results.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (739d8ef)
-
dataset: add LIBERO i2v/v2i robot manipulation retrieval (MOEB) (#5255)
-
dataset: add LIBERO i2v/v2i robot manipulation retrieval (MOEB)
Adds LIBEROI2VRetrieval and LIBEROV2IRetrieval, image<->video retrieval
over the LIBERO robot manipulation benchmark (40 tasks, 1,693 episodes,
256x256 @ 10 fps). Queries are goal-state images (final frames) of
held-out episodes; relevance is task-level and multi-positive with the
query's source episode excluded from the corpus, so exact frame matching
cannot solve the task. Adds the v2i task category (previously empty
direction) and a Robotics task domain.
Reference runs (ndcg@10): mteb/baseline-random-encoder 0.0297 (i2v) /
0.0261 (v2i); jinaai/jina-embeddings-v5-omni-nano 0.3537 (i2v) /
0.3312 (v2i).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
- fix: clarify LIBERO relevance definition and correct license to cc-by-4.0
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (718d295)
-
Add Clotho moment dataset (AT2A) (#5287)
-
Add clotho-moment retrieval task
-
update version
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
- update bibtex title
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
- Add task description and fix tests
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
-
Move descriptive stage to right directory.
-
Fix formatting in Clotho moment retrieval file
-
Update test_task_quality.py
-
Update test_task_quality.py
Co-authored-by: nehal309 <20415842+nehal309@users.noreply.github.com> (bb05784)
-
dataset: add EDIR (#5247)
-
dataset: add EDIR
-
update bibtex format and mark known duplicates
-
add script for processing data
-
lint format
-
add prompt
-
update query prompt and delete document prompt (
2440f68) -
Update GreenLeaf Law Embed Tiny: add 35+ languages (#5304)
Update GreenLeaf languages: add 35+ supported languages
Model trained on multilingual legal corpus covering 35+ languages
including English, German, French, Spanish, Chinese, Japanese,
Korean, Arabic, Hindi, and others.
Co-authored-by: Surya-saka <surya@judicialmind.ai> (626de72)
-
[MOEB] Add UniME-V2-LLaVA-OneVision-8B model (#5296)
-
Add UniME V2 model
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
- chore: remove tests
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
- Delete mteb_mock_run_results.md
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (9ef064e)