Skip to content

2.20.2

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 25 Aug 13:45
· 191 commits to main since this release

2.20.2 (2026-08-25)

Ci

  • ci: fix dataset_loading workflow hang by replacing uv sync with uv pip install (#5290)

  • ci: isolate uv cache to diagnose dataset_loading hang

  • ci: trigger dataset_loading workflow on changes to itself

  • ci: add venv creation diagnostics and restore pre-created venv workaround

  • ci: install Python before diagnostic step

  • ci: add trace logging and --no-install-project isolation test

  • ci: scope RUST_LOG to uv= to reduce trace volume

  • ci: replace uv sync with pip install to fix hang

uv sync hangs indefinitely because it must parse uv.lock (54 MB, 19k lines)
before doing anything else. Rust serde deserialization of this file likely
exhausts the runner's available memory, causing the process to stall.
The dataset-loading test only needs mteb, pytest, and pytest-rerunfailures,
so we bypass the lockfile entirely with a direct pip install.

  • ci: use uv pip install instead of uv sync to avoid lockfile parsing

  • ci: add --system flag to uv pip install

  • ci: create venv via python -m venv then uv pip install

  • ci: add bibtexparser dependency

  • ci: add gitpython dependency (8df869d)

  • ci: use minimal deps and fallback venv for dataset_loading workflow (#5289) (f5f504b)

  • ci: use Python 3.10 in dataset_loading workflow (#5286)

  • ci: use Python 3.10 in dataset_loading workflow to share uv cache with lint

  • ci: add 30m timeout and verbose output to dataset_loading install step (526cd73)

  • ci: use --frozen flag in dataset_loading and lint install steps (#5277)

  • ci: use --frozen flag in dataset_loading and lint install steps

Bypasses dependency resolution in CI where the lockfile should always
be used as-is, speeding up the install step.

  • ci: use --frozen flag in documentation install step (ea5bbe6)

Fix

  • fix: Enable PT ruff rule (#5238)

  • enable PT ruff rule

  • PT018 is not enforced and PT006 default apply (d088296)

Unknown

  • Update invalid links zensical setting (#5301)

update invalid links setting (3bf8702)

  • Fix jinav4 expriments (#5095)

  • use device during compute

  • fix experiment kwarg pass

  • fix dense multimodal

  • change model type to list (1459881)

  • model: Add gve models (#4975)

  • model: add GVE video embedding models (3B, 7B)

Adds Alibaba-NLP/GVE-{3B,7B}, general video embedders built on
Qwen2.5-VL that support text, image, video, and composed queries.
The HF repos ship a custom Qwen25VLForEmbedding class, but it is a
plain subclass of Qwen2_5_VLForConditionalGeneration with no extra
weights (and its remote code is incompatible with transformers
>= 4.56), so the native class is loaded instead and embeddings are
read from output_hidden_states, skipping vocab logits via
logits_to_keep=1. Pooling follows the model card: L2-normalized
last-token hidden state with left padding. Video frame sampling
reuses FramesCollator (fps=1, max 8 frames, per the model card).

Verified on MPS: text, image, and video inputs each produce
normalized 2048-dim embeddings for GVE-3B.

Closes #3770

  • fix: cap GVE video pixel budget via processor size dict

transformers 5.0 video processors read the pixel budget from
size["longest_edge"] and ignore the legacy max_pixels attribute,
so videos were processed at full resolution and overflowed
max_length, truncating vision tokens. Set both fields.

  • fix: denser default frame sampling for GVE (fps=2, max 32 frames)

The 8-frame demo settings from the model card under-sample videos for
retrieval; MSRVTT R@1 came in well below the paper. fps=2 capped at 32
frames matches other mteb video wrappers while keeping video tokens
(~800) within the 1200 max_length budget.

  • fix: raise GVE max_length to 4096 for dense video batches

32-frame videos tokenize past the demo's 1200 max_length, truncating
vision tokens which the processor rejects. 4096 leaves headroom; text
batches are unaffected since padding is to longest-in-batch.

  • review: use apply_chat_template and call-time video kwargs
  • build prompts via processor.apply_chat_template (output verified
    byte-identical to the previous manual template, so scores are
    unaffected)
  • pass the video pixel budget per call via videos_kwargs instead of
    mutating video_processor attributes; note max_pixels is ignored at
    call time in transformers 5, size={shortest_edge, longest_edge} is
    the working equivalent (verified: 1203 -> 727 tokens on a test clip)
  • drop low_cpu_mem_usage and use_fast (defaults)
  • fix: pool GVE embeddings at the appended <|endoftext|> token

The released GVE checkpoints pool the <|endoftext|> token appended
after the assistant turn (the convention of the team's GME codebase),
not the bare generation prompt shown in the model card demo. Without
it, retrieval quality drops sharply and task instructions actively
hurt. Verified on the authors' own UVRB MSRVTT split (1,000 JSFusion
pairs, 8 uniform frames, 200 tokens/frame): R@1 improves from 0.312
to 0.440 vs the paper's reported 0.431, and instruction sensitivity
collapses to noise (all placements 0.424-0.440).

  • lint: explicit strict=True in encode batch zip (a233646)

  • [MOEB] model: add ViCLIP video-language model (L-14, B-16) (#5013)

  • model: add ViCLIP video-language model (L-14, B-16)

    Adds ViCLIPWrapper and ModelMeta for OpenGVLab/ViCLIP-L-14-hf and
    OpenGVLab/ViCLIP-B-16-hf. Model from InternVid (ICLR 2024,
    arXiv:2307.06942). Covers T2V and V2T retrieval tasks in MOEB.
    Closes #5012, part of #4842.

  • style: fix ruff formatting in viclip_models.py

  • fix lint

  • fix: add mean/std source comment, remove defensive tensor checks


Co-authored-by: Rakshitha Ireddi <rakshithaireddi@Rakshithas-MacBook-Pro.local>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (0045490)

  • task: add ABOI2VRetrieval image-to-video retrieval (i2v) (#5284)

  • Add ABOI2VRetrieval: natively cross-modal image-to-video retrieval

Amazon Berkeley Objects ships, for the same product, a 360-degree turntable
"spin" photographed in a rig and separate catalog photographs shot at a
different time under different lighting. A query is therefore never a frame of
its own positive video and no crop, re-encode or temporal-neighbour
relationship links the two, so frame leakage is structurally impossible rather
than filtered out after the fact.

corpus 2857 videos, one per product, encoded from that product's spin by
selecting azimuth % 3 == 0. That single rule yields exactly 24 frames at
15-degree steps for all 8209 ABO spins (the 8116 dense ones store
azimuths 0..71, the 93 sparse ones store exactly {0,3,...,69}), so the
corpus is homogeneous without dropping any sequence. h264 / 384px long
side / 12 fps / 2.0 s, yuv420p limited-range bt709. Every output was
ffprobe-validated to decode to 24 frames.
queries 2857 catalog photographs, one per product, from the "context" bucket:
the product shot in a room or scene. Images sharing a perceptual hash
with another product (boilerplate, size charts), failing a zero-shot
category gate (swatches, macro crops, dimension diagrams), within a
pHash radius of the product's own spin, or shot on a white studio sweep
are all excluded.
qrels 1:1, score 1. One product per spin sequence, so no two queries share a
relevant document.

Queries are selected by a semantic category rather than by distance from the
answer, so every query stays an answerable depiction of the item; difficulty
comes from corpus size and the domain gap between a styled room photo and a
turntable render. Product types are restricted to five volumetric home-goods
categories; flat goods (RUG, WALL_ART) are excluded because a turntable
rotation of a flat object is close to degenerate.

ABO is CC BY 4.0, which permits redistribution of adapted material with
attribution. Note the bucket still carries a stale LICENSE-CC-BY-NC-4.0.txt
from 2021 and the AWS open-data registry entry was never updated after the
2023 relicense; the current README, all four subdirectory READMEs and the
project page all state CC BY 4.0.

i2v only: "v2i" is not in the TaskCategory literal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

  • Pin ABO-I2V to the revision carrying the CC BY 4.0 dataset card

The initial push produced only the auto-generated card, which has no license
or attribution. CC BY 4.0 Section 3(a) requires the creator credit, license
notice, warranty disclaimer, source link and an indication that the material
was modified to be present where the material is shared, so the card now
carries all of it plus the note about the bucket's stale NC license file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

  • Shorten ABOI2VRetrieval description

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> (429b5a6)

  • Add Qwen3 voice models (#5240)

  • Add Qwen3 voice models

Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>

  • update audio collator and memory

Co-authored-by: nehal309 <20415842+nehal309@users.noreply.github.com> (33345dc)

  • dataset: add MovingFashion bidirectional image-video retrieval (#5272)

  • dataset: add MovingFashion video-to-image retrieval

  • dataset: add reverse MovingFashion retrieval (d1fb15f)

  • Add judicialmind/greenleaf-law-embed-tiny (#5274)

  • Add judicialmind/greenleaf-law-embed-tiny model implementation

  • 596M-parameter legal-domain embedding model
  • Bidirectional attention on Qwen3 architecture
  • 1024-dim embeddings, 32K context, mean pooling
  • Built-in int8/binary quantization via custom code
  • MTEB(Law, v1): 64.49% mean NDCG@10 (8/8 tasks)
  • MLEB-12: 78.41% mean NDCG@10
  • Apache 2.0, open weights
  • Custom wrapper handles trust_remote_code loading
  • Training data: proprietary (not disclosed)
  • Add official mteb mock-run results (28/28 passed)

  • Update revision hash to match scrubbed model repo (bff06c94)

  • Remove custom GreenLeafEmbedWrapper, use SentenceTransformerEncoderWrapper directly with trust_remote_code kwarg

  • Remove mteb_mock_run_results.md

  • Add public_training_data link to judicialmind/legal-training-dataset

  • Update mteb/models/model_implementations/greenleaf_models.py


Co-authored-by: Surya-saka <surya@judicialmind.ai>
Co-authored-by: Surya-saka <sakasurya@users.noreply.huggingface.co>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (d1665ae)

  • task: Add Breakfast Clustering and Breakfast Pair Classification datasets (#5275)

  • [MOEB] Add Breakfast Clustering and Breakfast Pair Classification datasets

  • remove var _BIBTEX

  • add descriptive stats

  • move data to mteb (2c1ceab)

  • dataset: Add EMID A2I + I2A retrieval datasets (#5279)

  • [MOEB] Add EMID pair classification and EMID a2i + i2a retreival datasets

  • add stats

  • fix types

  • reduce size and dedup

  • remove pc files

  • move data to mteb org (43ed975)

  • Add Webvid covr dataset (#5216)

  • Add WebVid dataset and retrieval task

WebVid provides Video1 + Edit = Video2 examples. Rather than using Video1, they have proposed to use middle frame of the video as Image, which makes the task Image + Test -> Video retrieval task.

Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>

  • Exclude corrupted video

  • Add descriptive stats

  • fix revision and arxiv url

  • fix reference and main_score

  • Add WebVid dataset and retrieval task

WebVid provides Video1 + Edit = Video2 examples. Rather than using Video1, they have proposed to use middle frame of the video as Image, which makes the task Image + Test -> Video retrieval task.

Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>

  • Exclude corrupted video

  • Add descriptive stats

  • fix revision and arxiv url

  • fix reference and main_score

  • add WebVidCoVRIT2VRetrieval to known duplicate exception


Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>
Co-authored-by: Deep Shah <shahdeep@google.com> (aeb24c4)

  • task: add EVVE event video retrieval (v2v) (#5197)

  • dataset: add EVVE retrieval

  • dataset: align EVVE with review conventions

  • dataset: finalize EVVE review alignment

  • fix: download EVVE construction metadata

  • fix: pin updated EVVE dataset revision (0c0f6e5)

  • Remove AfriMTEB task dataset_transform() (#4905)

  • Fix AfriMTEB task schema issues

  • Remove unrelated benchmark changes

  • Remove unnecessary AfriXNLI dataset transform

  • Remove unrelated benchmark changes

  • Remove commented out code in AfriHate and KinNews classification

  • Remove dataset_transform from AfriHate and KinNews classification


Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (0fb6484)

  • dataset: Add multilingual MMarco retrieval task (#5068)

  • feat: add multilingual MMarco retrieval task

  • added revision MMarcoRetrievalMultilingual

  • Trigger CI after updating PR description

  • added MMarcoRetrievalMultilingual and stats

  • added MMarcoRetrievalMultilingual to mteb/tasks/retrieval/multilingual/init.py

  • changed task description

  • Marked MMarcoRetrieval as superseded by the new MMarcoRetrievalMultilingual task

  • MMarcoRetrievalMultilingual task contributed_by=None

  • removed contributed_by and is_beta

  • Add MMarcoRetrievalMultilingual to KNOWN_ISSUES in test_task_quality

  • Add MMarcoRetrievalMultilingual to KNOWN_ISSUES short_text


Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (ad874e1)

  • dataset: Add SciRepEval classification tasks (DRSM, biomimicry, FoS, MeSH) (#5026)

  • Add SciRepEval DRSM classification task

Add SciRepEvalDRSMClassification, a 5-way single-label text
classification task from the SciRepEval benchmark (Singh et al., EMNLP
2023). Given a biomedical paper's title and abstract, the task predicts
its Disease Research State Model category.

The allenai/scirepeval drsm config exposes only a single evaluation
split, so dataset_transform builds a stratified train/test split for the
classification evaluator and subsamples the test split to 2048 examples.

Includes TaskMetadata (revision-pinned), registration in the eng
classification init, and precomputed descriptive statistics.

Part of #591.

  • Add more SciRepEval classification tasks (biomimicry, FoS, MeSH)

Extend the SciRepEval coverage beyond DRSM with three more
classification tasks from the allenai/scirepeval benchmark:

  • SciRepEvalBiomimicryClassification: binary relevance classification
    (AbsTaskClassification). Single "evaluation" split -> stratified
    train/test split.
  • SciRepEvalFoSClassification: multi-label Fields of Study
    classification (AbsTaskMultilabelClassification), predicting MAG
    fields such as Chemistry / Materials science.
  • SciRepEvalMeSHDescriptorsClassification: 30-class MeSH descriptor
    classification (AbsTaskClassification). Rows are (paper, descriptor)
    pairs, so papers are deduplicated before the train/test split to
    avoid a single abstract leaking across splits.

For FoS and MeSH only the "evaluation" split is loaded (the upstream
train splits are hundreds of thousands to millions of rows). Each task
ships pinned metadata, registration, and precomputed descriptive
statistics.

Reference runs (accuracy, CPU): random-encoder vs multilingual-e5-small

  • Biomimicry: 0.510 -> 0.728
  • FoS: 0.001 -> 0.178
  • MeSH: 0.033 -> 0.582

Part of #591.

  • Fix CI: dedup DRSM duplicate text + allowlist FoS for typos
  • test_dataset_quality flagged one duplicate title+abstract in the DRSM
    train split (6169 samples / 6168 unique); filter duplicate texts before
    the train/test split so no document leaks across the split. Regenerated
    descriptive stats accordingly.
  • typos flagged the FoS (Fields of Study) acronym in the multilabel task;
    add SciRepEvalFoSClassification and FoS to the typos identifier allowlist.
  • Address review: reframe task descriptions + set license/annotations_creators

Set license to odc-by for the four SciRepEval tasks (aggregate benchmark
is released under ODC-BY per the SciRepEval repo). Set biomimicry
annotations_creators to human-annotated (labels are manually annotated
gold tags from the PeTaL database). Rewrite the four task descriptions in
a what-it-tests / how / attributes style.

  • Switch single-label SciRepEval tasks to cross-validation

  • Keep evaluation split name instead of renaming to train

Per Samoed's review comment, the cross-validation SciRepEval tasks
(DRSM, biomimicry, MeSH descriptors) no longer rename their single
split to train. eval_splits and train_split are both set to
evaluation, and the descriptive stats files are updated to match.


Co-authored-by: Claude <noreply@anthropic.com> (14b295c)

  • benchmark: Add Slovak tasks and SK-MTEB benchmark (MTEB(slk, v1)) (#4788)

  • Add Slovak tasks and SK-MTEB benchmark (MTEB(slk, v1))

Adds new Slovak-language tasks across 7 task types and registers the full SK-MTEB benchmark. Tasks cover retrieval, STS, pair classification, classification, reranking, clustering, and bitext mining for Slovak.

  • fix: Update Slovak citations and improve dataset references

  • Update mteb/benchmarks/benchmarks/benchmarks.py

Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>

  • Update mteb/tasks/classification/slk/multi_eup_slovak_classification.py

Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>

  • fix: Update citation for SkMTEB benchmark to reflect new publication details

  • refactor: Simplify Multi-EuP Slovak classification tasks by removing custom dataset loader and mixin

  • fix: Update SkMTEB tasks after resolving conflicts and update reference with arXiv link

  • Fix SlovakRTE label polarity

  • fix: Update dataset_transform signature

  • Enable trust_remote_code for gte-multilingual-base

Alibaba-NLP/gte-multilingual-base ships its architecture as custom
code in the HF repo (Alibaba-NLP/new-impl); without this it fails to
load. Mirrors the existing pattern in arctic_models.py.

  • fix: Improve dataset descriptions for Slovak NLI, RTE, SkQuadReranking, SMESumRetrieval, and STS

  • fix: Update method signatures across Slovak tasks to standardize dataset_transform and load_data definitions

  • refactor: Point SlovakSumURLClustering, SlovakSTS, SMESumRetrieval at pre-built mteb datasets

  • fix: Exempt SlovakPharmacyDrMaxReranking and OpusSlovakEnglishBitextMining from dataset quality checks

Both fail new duplicate/short-text checks due to genuine, negligible source-data
characteristics rather than bugs: OPUS-100 naturally repeats short common
subtitle/legal-document phrases (24/2000 test pairs), and DrMax's real search-query
log contains a couple of 1-character queries (2/4676).


Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (c8a519a)

  • model: add Qwen3-VL-Reranker (2B, 8B) (#5057)

  • model: add Qwen3-VL-Reranker (2B, 8B)

  • fill languages from model card (33 languages)

  • trigger CI re-run

  • reranker: add use_instructions flag instead of mutating kwargs


Co-authored-by: Hubert Lu <hubielu@email.com> (233478f)