Skip to content

2.18.17

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 13 Aug 12:28
· 277 commits to main since this release

2.18.17 (2026-08-13)

Ci

  • ci: Fix Copilot review comment command (#5172)

Fix Copilot review comment command

Updated the comment command to use '@copilot' instead of '@github-copilot'. (04403ce)

  • ci: add Copilot dataset PR review via label trigger (#5163)

  • add Copilot review instructions for dataset PRs

  • ci: trigger Copilot review via label instead of applyTo

  • ci: clarify results table and add random encoder run command

  • ci: remove unverifiable dataset runs check, add gap and size checks (728e9cb)

Fix

  • fix: require datasets>=4.0.0 for video extra (#5173)

  • fix: require datasets>=4.0.0 for video extra

  • regenerate uv.lock (1d587bf)

  • fix: combine subsets in run_settings.jsonl (#5099)

  • fix combine subsets in run_settings.jsonl

  • copilot suggestion

  • matched implementation with existing run_settings

  • fixed typecheck (c6b7e1b)

Unknown

  • Enable some ruff rules (#5108) (c4f9d80)

  • model: add PS3 (nvidia/PS3-1.5K/4K-SigLIP and SigLIP2) (#5077)

Co-authored-by: Hubert Lu <hubielu@email.com> (df925fd)

  • Add Flickr dataset I2A and A2I (#5137)

  • Add Flickr dataset I2A and A2I

Co-Authored-By: Deep Shah <21212684+deep9539@users.noreply.github.com>

  • refactor and directory update

    1. Add descriptive stats and 2. HF commit version when reading
  • format lint files


Co-authored-by: Deep Shah <21212684+deep9539@users.noreply.github.com> (0e07d11)

  • dataset: add CoVR-R (#5116)

  • dataset: add CoVR-R

  • reupload dataset and update descriptive stats (cce148b)

  • dataset: add REAL-MM-RAG (#5106)

  • dataset: add REAL-MM-RAG

  • test: allow duplicate images in REAL-MM-RAG

  • Update mteb/tasks/retrieval/eng/real_mm_rag_retrieval.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Update mteb/tasks/retrieval/eng/real_mm_rag_retrieval.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (4b1fbb0)

  • task: add Stanford I2V image-to-video+audio retrieval (i2va) (#5151)

  • dataset: add Stanford I2V retrieval

  • dataset: add Stanford I2V audio

  • dataset: align Stanford I2V license metadata (8e170e9)

  • model: add Dasheng audio encoders (base, 0.6B, 1.2B) (#5118)

  • wip: Dasheng audio wrapper

  • model: add Dasheng audio encoders (base, 0.6B, 1.2B)

  • deps: declare einops for Dasheng

  • revert unrelated uv.lock changes

  • meta: add citation and training datasets for Dasheng


Co-authored-by: Hubert Lu <hubielu@email.com> (cf631a3)

  • model: add Cosmos-Embed1 (nvidia/Cosmos-Embed1-224p/336p/448p) (#5133)

  • model: add Cosmos-Embed1 (nvidia/Cosmos-Embed1-224p/336p/448p)

  • model: add Cosmos-Embed1 (nvidia/Cosmos-Embed1-224p/336p/448p)

  • fix: GPU-path bugs in Cosmos-Embed1 (int device, dtype cast, mixed-resolution clips)

  • review: access projections directly instead of a getattr helper


Co-authored-by: Hubert Lu <hubielu@email.com> (7525053)

  • model: add NVIDIA RADIO family (RADIO-B/L/H) (#5061)

  • model: add NVIDIA RADIO family (RADIO-B, RADIO-L, RADIO-H)

  • fix: cite RADIOv2.5 paper alongside AM-RADIO

  • Apply suggestions from code review

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • fix: lint after applying review suggestions

  • chore: refresh mock-run results

  • refactor: drop redundant modality guards, remove mock-run results

  • Update mteb/models/model_implementations/radio_models.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (d35c86e)

  • dataset: add SpokenCOCO (#5121) (f2dfbd3)

  • task: add CaReBench video retrieval (#5112)

  • task: add CaReBench video retrieval

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

  • task: pin CaReBench dataset and add descriptive stats

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> (e040e04)

  • Add Audio flamingo 3 (#5080)

  • Add audio flamingo 3 model

  • Fix type errors

Fixing type error and minor refactoring.

Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>

  • fix the cast error using chat template

fix the cast error using chat template
Solves expected scalar type Float but found BFloat16

Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>.

  • Add auto cast

Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>

  • Convert embedding to float32 before returning

embeddings was bfloat16 (produced under torch.autocast(dtype=torch.bfloat16)), and NumPy has no bfloat16 dtype, so the final torch.cat(...).numpy() call raised TypeError: Got unsupported ScalarType BFloat16. Casting to float32 before moving off-GPU fixes it — only the small pooled embedding tensor is upcast, not the full model.

Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>

  • Update mteb/models/model_implementations/audio_flamingo.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Minor nits

  • Change init func signature to capture torch_dtype and device_map

  • Add n embedding params

  • fix lint

  • fix lint

  • fix import order

  • Apply suggestions from code review

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>


Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (09e0c88)

  • dataset: add MMLongBench-Doc (#5138)

  • dataset: add MMLongBench-Doc

  • test: register MMLongBenchDocRetrieval duplicate images (7bd35b3)

  • model: add malteos/most-embed-de (German retrieval, Nemotron-3-Embed-1B fine-tune) (#5149)

Adds a ModelMeta for malteos/most-embed-de, a 1.1B German retrieval embedding
model fine-tuned from nvidia/Nemotron-3-Embed-1B-BF16.

The base model is already implemented, so this reuses SentenceTransformerEncoderWrapper
and the existing nemotron-3-embed extras group. The 4096 model_max_length cap mirrors
the base model's entry so results stay comparable to the baseline.

Verified with mteb mock-run (28/28 text tasks passing).

Claude-Session: https://claude.ai/code/session_01LaEay2SrYBRYnojLT8mJi4 (5e36d33)

  • Add KiteFishAI/Nano-Em1-0.6B-v2.1 (#5150)

  • Add KiteFishAI/Nano-Em1-0.6B-v2.1

  • Add new training datasets to kite_fish_models.py

  • Refactor training datasets to use dictionary format (8df42c1)

  • dataset: add Greatest Hits audio-video material retrieval (a2v, v2a) (#5003)

  • task: add Greatest Hits audio-video material retrieval (a2v, v2a)

Impact-sound <-> video retrieval from Greatest Hits / Visually Indicated Sounds
(CC-BY-4.0). Match an impact across modalities by material (17 materials). 992
impacts. LCO-Omni: a2v 15.2 / v2a 17.4 vs random ~10.5 nDCG@10 (map@10 3x random).
Includes descriptive statistics.

  • fix: embed GreatestHits audio into the dataset (portability)

The audio side was stored as file-path references (via Dataset.to_parquet
with an Audio() column), so it only decoded on the original build machine.
Rebuilt with the audio bytes embedded directly in the parquet (matching the
video side); refreshed to 1000 clips over 17 materials. Updated dataset
revisions and regenerated descriptive statistics. (e73b747)

  • Remove gradio lb tests and ci (#5103)

  • remove gradio lb tests and ci

  • remove duplicated import (409a355)

  • task: add ADVANCE audio-image retrieval (a2i, i2a) (#5119)

  • task: add ADVANCE audio-image retrieval (a2i, i2a)

Adds ADVANCEA2IRetrieval and ADVANCEI2ARetrieval, filling the audio<->image
gap under MOEB Track 1 (#4842). 5,075 geotagged locations pairing FreeSound
field recordings with co-located Google Earth aerial imagery across 13
land-cover classes, license CC-BY-4.0 (verified via the dataset's Zenodo
record, the primary source repo). Reshaped from blanchon/ADVANCE into
BEIR-style corpus/queries/qrels layout, pushed to yaswanth169/ADVANCE-A2I
and yaswanth169/ADVANCE-I2A. Note: ~6.7GB at original resolution.

Closes #5006

  • docs: add ADVANCE dataset construction script to the PR

Per Samoed's review on #5119 -- the script that reshapes blanchon/ADVANCE
into BEIR-style corpus/queries/qrels layout was only linked in a PR
comment; now committed under scripts/ as part of the diff itself. (a3feb86)

  • model: add ColQwen-Omni (vidore/colqwen-omni-v0.1) (#5115)

  • model: add ColQwen-Omni (vidore/colqwen-omni-v0.1)

  • fix: add n_embedding_parameters to ColQwen-Omni meta

  • fix: set n_embedding_parameters for ColQwen-Omni

  • revert unrelated uv.lock changes

  • fix: set do_sample_frames=False once at init for ColQwen-Omni


Co-authored-by: Hubert Lu <hubielu@email.com> (3c1f36d)

  • add bright pro benchmark (#5129) (44e33bf)

  • model: add Singaraj/morisien-embed (Mauritian Creole, mfe) (#5125) (1276f64)

  • benchmark: add BRIGHT-Pro retrieval (7 StackExchange domains) (#4929)

  • task: add Bright-Pro retrieval (7 StackExchange domains)

Adds seven per-domain retrieval tasks built on yale-nlp/Bright-Pro:
biology, earth_science, economics, psychology, robotics, stackoverflow,
sustainable_living. Each task loads the documents and examples HF
configs and exposes binary qrels from gold_ids for standard nDCG@10
evaluation. The dataset's reasoning-aspect annotations (aspects config)
are not consumed by the standard retrieval task and remain available on
the Hub for users who want aspect-aware evaluation.

Closes #4623

  • benchmark: register BRIGHT-Pro as a Benchmark

Register a top-level BRIGHT_PRO Benchmark grouping the 7 per-domain
BrightPro* retrieval tasks so users can run mteb.get_benchmark(&#34;BRIGHT-Pro&#34;)
in one call. Mirrors how BRIGHT/BRIGHT(v1.1) are registered.

Address review comment from @Samoed on #4651.

  • task: align BrightPro prompts with BRIGHT-Pro paper convention

Update each BrightPro{Domain}Retrieval task's query prompt from the
BRIGHT-v1.1-style 'Represent this {domain} post for searching relevant
passages: ' to the BRIGHT-Pro paper's 'Given a {domain} post, retrieve
relevant passages that help answer the post'.

Matches the prompt body used by BRIGHT-Pro's reference evaluation harness
across all instruction-tuned retrievers (qwen3-embed, reasonir, gte-Qwen2,
gritlm, etc.) so MTEB users running these models on BrightPro tasks see
the paper-protocol numbers.

  • stats: add descriptive statistics for the 7 BrightPro retrieval tasks

Generated via task.calculate_descriptive_statistics() as required by
tests/test_tasks/test_metadata.py; fixes the failing test CI jobs.

  • model: add BrightPro prompts to ReasonIR prompts_dict

Mirror the existing per-task BRIGHT entries with the 7 BrightPro tasks,
verbatim-matching the task metadata prompts so scores reproduce out of
the box (requested in PR #4651 review).

  • reupload

  • task: use natural domain names in BrightPro prompts

The raw subset slugs (earth_science, sustainable_living, stackoverflow)
leaked into the query prompts because BRIGHT's config template
substitutes the subset key directly (Given a {task} post, ...), and
BRIGHT-Pro inherited that mechanism. Replace them with the StackExchange
site names they refer to, and fix the article agreement (a -> an) for
Earth Science / Economics.

Keeps mteb/models/model_implementations/reasonir_model.py in sync.

  • docs: lead task/benchmark descriptions with what they measure

The leaderboard renders the first lines of a description, so state the
retrieval quality being measured before the details. Move the benchmark
attribution to the end and rephrase it as 'Was developed as part of',
since a task can belong to several benchmarks.

Also set the benchmark display_name to BRIGHT-Pro.

  • Revert "task: use natural domain names in BrightPro prompts"

This reverts commit b021cef.

@Samoed is right that rewriting the prompts shifts the scores, which would
leave the leaderboard numbers not matching the ones published with the
benchmark. Keeping the prompts as the paper ran them takes priority over the
cosmetics of the subset slugs, so this goes back to the original wording.

The slug leakage (a "sustainable_living post", "a earth_science post") is real
but has to be weighed against reproducibility; see the PR thread for the
measured effect.

  • Reapply "task: use natural domain names in BrightPro prompts"

This reverts commit 4bc04e6.


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (f0cb9fd)

  • dataset: add MorisienMTBitextMining (Mauritian Creole, mfe) (#5110)

  • dataset: add MorisienMTBitextMining (Mauritian Creole, mfe)

  • dataset: set MorisienMTBitextMining license to MIT

The upstream MorisienMT dataset has been relicensed to MIT by its author, so the
task metadata and the repackaged dataset now declare MIT.

  • dataset: repoint MorisienMTBitextMining to mteb-hosted copy (a0cf754)

  • model: add erikkaum/lattice-retrieval (#5105)

  • feat: add lattice retrieval model

  • test: add lattice mock-run results

  • Delete mteb_mock_run_results.md


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (98f3b8f)

  • dataset: add LPMusicCapsMTT A2T and T2A retrieval tasks (#5071)

  • dataset: add LPMusicCapsMTT A2T and T2A retrieval tasks

  • dataset: add creation script for LPMusicCapsMTT


Co-authored-by: Hubert Lu <hubielu@email.com> (a57299e)

  • model: moca-embed/MoCa-Qwen25VL-3B (#3589) (#5076)

  • feat: add MoCa-Qwen25VL-3B (closes #3589)

  • fix: add n_embedding_parameters for MoCa

  • fix: address review on MoCa (default prompt, training datasets, transformers pin)

  • chore: regenerate uv.lock for the moca extra


Co-authored-by: Hubert Lu <hubielu@email.com> (09a4a30)

  • model: add cnmoro/static-nomic-384-pten-v2-st (static pt/en Model2Vec) (#5075)

  • model: add cnmoro/static-nomic-384-pten-v2

Static (Model2Vec/Tokenlearn) pt/en embedding model distilled from
nomic-embed-text-v2-moe. Adds Model2VecStaticModelWrapper because the
checkpoint uses vocabulary quantization, which sentence-transformers'
StaticEmbedding cannot load.

  • address review: reuse Model2VecModel, declare model2vec extra, drop citation
  • reuse the existing Model2VecModel loader instead of a duplicate wrapper
  • add extra_requirements_groups=["model2vec"] for the required dependency
  • citation=None (the model uses the technique; it should not cite that work)
  • drop the borrowed MODEL2VEC_CITATION block and the placeholder docstring
  • framework: drop "Sentence Transformers" (this checkpoint cannot load under it)
  • switch to sentence-transformers export static-nomic-384-pten-v2-st

Publishes a materialized (one row per token) export of the model so the standard
SentenceTransformerEncoderWrapper can load it, instead of relying on the deprecated
Model2VecModel. All 186 tasks re-run against the new revision. (ed285c7)

  • model: Declare Common Voice in fusion-embedding training_datasets (contamination disclosure) (#5097)

Declare Common Voice (and derived CommonLanguage) in fusion-embedding training_datasets (0204e96)