Skip to content

2.20.1

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 23 Aug 08:33
· 217 commits to main since this release

2.20.1 (2026-08-23)

Fix

Unknown

  • model: Add LanguageBind video and audio model wrapper (#4557)

  • model: add LanguageBind video, audio, and image wrappers with compatibility patches

  • feat: HF mirror auto-download for LanguageBind + extras group for reproducibility

  • wip: phase 1 review fixes for GPU testing

  • refactor: trim LanguageBind compat shims to the two that are required

Tested each of the five compatibility patches individually on GPU against the
pinned transformers 4.45.2 / torchvision 0.20.1 / torchaudio 2.5.1:

  • clip_loss: transformers already defines it; the hasattr guard never fires -> removed
  • torchaudio.set_audio_backend: still present as a no-op; guard never fires -> removed
  • _expand_mask: LanguageBind hard-imports it from modeling_clip -> kept
  • torchvision.transforms.functional_tensor: pytorchvideo hard-imports it -> kept
  • tokenizer init shim: kept defensively (not triggered under the pinned
    versions, but guards against transformers load-path drift)

Added inline comments documenting why the two retained shims are required.

  • refactor: apply compat shims in the loader instead of at module import

Move the _patch_compatibility() call into _ensure_languagebind_source() so the
shims only run when a LanguageBind model is actually loaded, rather than
globally at module import time.

  • refactor: move LanguageBind compat fixes into the pinned source snapshot

Instead of monkey-patching transformers/torchvision/torchaudio at runtime from
the MTEB wrapper, the compatibility fixes now live in the vendored LanguageBind
source itself (pinned HF revision d2d0f6f):

  • _expand_mask / clip_loss: vendored into the modeling files (removed from
    transformers in the attention-mask refactor, before 4.40)
  • tokenizer init: pass args to CLIPTokenizer.init by keyword
  • torchvision.transforms.functional_tensor: aliased before importing pytorchvideo
  • torchaudio.set_audio_backend: guarded with hasattr

The wrapper now only downloads and imports the pinned source revision, with no
runtime patching of external libraries.

Verified clean import under torch 2.11 / torchvision 0.26 / transformers 4.45.2.

  • chore: slim languagebind extra and drop pycache from pinned source
  • Remove huggingface_hub (already a core dependency), librosa and soundfile
    (LanguageBind doesn't import them; audio dataset decoding uses the existing
    audio extras) from the languagebind extra
  • Bump pinned LanguageBind source revision to a226c49b (drops pycache)
  • chore: pin languagebind extra dependencies

Add lower bounds to the languagebind extra deps for reproducibility:
einops, decord, opencv-python-headless, pytorchvideo, peft.

  • chore: pin transformers in languagebind extra and add conflict entry
  • Cap transformers at >=4.40.0,<5.0.0 (LanguageBind targets 4.x; verified on 4.45.2)
  • Add torchvision/torchaudio to the extra (imported by the LanguageBind source,
    not in core)
  • Add languagebind to [tool.uv].conflicts since it pins transformers <5 while
    some other extras require ==5.0.0
  • refactor to use python package and added temp tests

  • fix: functional_tensor shim for modern torchvision + pin transformers<4.46 + use get_model in tests

  • Delete tests/test_models/test_language_bind.py

  • fix languagebind video transform and add missing audio/video deps

  • restore PLW0717 ruff ignore rule


Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (aa46f13)

  • add missing qrel ID statistics (#5226)

  • feat: add missing qrel ID statistics

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

  • fix: update retrieval statistics with missing qrel IDs

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

  • refactor: drop redundant dangling qrel comment per review

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

  • test: check qrels for missing query and corpus IDs

Use exact num_missing_query_ids / num_missing_corpus_ids statistics
instead of inferring dangling qrels from
unique_relevant_docs &gt; num_documents. Keep the old bound as a fallback
for statistics generated before these fields existed.

Regenerate OVENIT2ITRetrieval statistics and allowlist the four tasks
flagged by the new check. See PR description for details.


Co-authored-by: Xu Liu <lxer@Xus-MacBook-Pro.local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (d5a4595)

  • model: add V-SPLADE (#5265) (a905523)

  • dataset: add IntPhys2 video zero-shot task (#5225)

  • dataset: add IntPhys2 video zero-shot task

  • Move IntPhys2 constants into task class

  • dataset: use MTEB IntPhys2 artifact

  • dataset: add IntPhys2 conversion script (436d67f)

  • dataset: add Beehive States audio classification task (#5262)

  • dataset: add Beehive States audio classification task

  • fix: remove unnecessary writer batch setting

  • Update mteb/abstasks/classification.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • test: remove unnecessary classification sampler test

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (768fac5)

  • dataset: add RT-1 t2v/v2t robot manipulation retrieval (MOEB) (#5260)

Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3

Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (4e4b5b1)

  • add giga models (#5252)

  • add giga models

  • Update mteb/models/model_implementations/giga_models.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Update mteb/models/model_implementations/giga_models.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Update giga_models.py

added instructions, changes model loader to InstructSentenceTransformerModel

  • lint and remove extra fields

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4ca789c)

  • Add model: OpenGVLab/InternVideo2-CLIP-1B-224p-f8 (#5078)

  • feat: add InternVideo2-CLIP-1B-224p-f8 (WIP)

  • chore: pin InternVideo2 CLIP revision

  • fix: disable LoRA wrapper so InternVL-C text weights load

  • feat: add InternVideo2-CLIP-1B-224p-f8

  • disable LoRA wrapper so InternVL-C text weights load under peft>=0.7
  • stub flash-attn fused kernels; use the non-fused equivalent path
  • shim transformers symbols moved to pytorch_utils
  • MSVD t2v/v2t ndcg@10 0.859/0.858, mock-run 37/37
  • refactor: drop dead peft remap, fix install hint, add n_embedding_parameters

  • refactor: drop dead peft remap, fix install hint, add n_embedding_parameters

  • chore: declare internvideo2 as a conflicting extra, regenerate uv.lock

  • refactor: load InternVideo2-CLIP via SentenceTransformers

  • style: fix InternVideo2 lint

  • Update mteb/models/model_implementations/internvideo2_models.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • chore: drop mock run results from the diff

Co-authored-by: Hubert Lu <hubielu@email.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (2902127)

  • task: add visual-only Stanford I2V retrieval (#5258)

  • task: add visual-only Stanford I2V retrieval

  • refactor: colocate Stanford I2V task variants (a3f7b44)