2.20.1
2.20.1 (2026-08-23)
Fix
Unknown
-
model: Add LanguageBind video and audio model wrapper (#4557)
-
model: add LanguageBind video, audio, and image wrappers with compatibility patches
-
feat: HF mirror auto-download for LanguageBind + extras group for reproducibility
-
wip: phase 1 review fixes for GPU testing
-
refactor: trim LanguageBind compat shims to the two that are required
Tested each of the five compatibility patches individually on GPU against the
pinned transformers 4.45.2 / torchvision 0.20.1 / torchaudio 2.5.1:
- clip_loss: transformers already defines it; the hasattr guard never fires -> removed
- torchaudio.set_audio_backend: still present as a no-op; guard never fires -> removed
- _expand_mask: LanguageBind hard-imports it from modeling_clip -> kept
- torchvision.transforms.functional_tensor: pytorchvideo hard-imports it -> kept
- tokenizer init shim: kept defensively (not triggered under the pinned
versions, but guards against transformers load-path drift)
Added inline comments documenting why the two retained shims are required.
- refactor: apply compat shims in the loader instead of at module import
Move the _patch_compatibility() call into _ensure_languagebind_source() so the
shims only run when a LanguageBind model is actually loaded, rather than
globally at module import time.
- refactor: move LanguageBind compat fixes into the pinned source snapshot
Instead of monkey-patching transformers/torchvision/torchaudio at runtime from
the MTEB wrapper, the compatibility fixes now live in the vendored LanguageBind
source itself (pinned HF revision d2d0f6f):
- _expand_mask / clip_loss: vendored into the modeling files (removed from
transformers in the attention-mask refactor, before 4.40) - tokenizer init: pass args to CLIPTokenizer.init by keyword
- torchvision.transforms.functional_tensor: aliased before importing pytorchvideo
- torchaudio.set_audio_backend: guarded with hasattr
The wrapper now only downloads and imports the pinned source revision, with no
runtime patching of external libraries.
Verified clean import under torch 2.11 / torchvision 0.26 / transformers 4.45.2.
- chore: slim languagebind extra and drop pycache from pinned source
- Remove huggingface_hub (already a core dependency), librosa and soundfile
(LanguageBind doesn't import them; audio dataset decoding uses the existing
audio extras) from the languagebind extra - Bump pinned LanguageBind source revision to a226c49b (drops pycache)
- chore: pin languagebind extra dependencies
Add lower bounds to the languagebind extra deps for reproducibility:
einops, decord, opencv-python-headless, pytorchvideo, peft.
- chore: pin transformers in languagebind extra and add conflict entry
- Cap transformers at >=4.40.0,<5.0.0 (LanguageBind targets 4.x; verified on 4.45.2)
- Add torchvision/torchaudio to the extra (imported by the LanguageBind source,
not in core) - Add languagebind to [tool.uv].conflicts since it pins transformers <5 while
some other extras require ==5.0.0
-
refactor to use python package and added temp tests
-
fix: functional_tensor shim for modern torchvision + pin transformers<4.46 + use get_model in tests
-
Delete tests/test_models/test_language_bind.py
-
fix languagebind video transform and add missing audio/video deps
-
restore PLW0717 ruff ignore rule
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (aa46f13)
-
add missing qrel ID statistics (#5226)
-
feat: add missing qrel ID statistics
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- fix: update retrieval statistics with missing qrel IDs
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- refactor: drop redundant dangling qrel comment per review
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- test: check qrels for missing query and corpus IDs
Use exact num_missing_query_ids / num_missing_corpus_ids statistics
instead of inferring dangling qrels from
unique_relevant_docs > num_documents. Keep the old bound as a fallback
for statistics generated before these fields existed.
Regenerate OVENIT2ITRetrieval statistics and allowlist the four tasks
flagged by the new check. See PR description for details.
Co-authored-by: Xu Liu <lxer@Xus-MacBook-Pro.local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (d5a4595)
-
dataset: add IntPhys2 video zero-shot task (#5225)
-
dataset: add IntPhys2 video zero-shot task
-
Move IntPhys2 constants into task class
-
dataset: use MTEB IntPhys2 artifact
-
dataset: add IntPhys2 conversion script (
436d67f) -
dataset: add Beehive States audio classification task (#5262)
-
dataset: add Beehive States audio classification task
-
fix: remove unnecessary writer batch setting
-
Update mteb/abstasks/classification.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
- test: remove unnecessary classification sampler test
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (768fac5)
- dataset: add RT-1 t2v/v2t robot manipulation retrieval (MOEB) (#5260)
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (4e4b5b1)
-
add giga models (#5252)
-
add giga models
-
Update mteb/models/model_implementations/giga_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
- Update mteb/models/model_implementations/giga_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
- Update giga_models.py
added instructions, changes model loader to InstructSentenceTransformerModel
- lint and remove extra fields
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4ca789c)
-
Add model: OpenGVLab/InternVideo2-CLIP-1B-224p-f8 (#5078)
-
feat: add InternVideo2-CLIP-1B-224p-f8 (WIP)
-
chore: pin InternVideo2 CLIP revision
-
fix: disable LoRA wrapper so InternVL-C text weights load
-
feat: add InternVideo2-CLIP-1B-224p-f8
- disable LoRA wrapper so InternVL-C text weights load under peft>=0.7
- stub flash-attn fused kernels; use the non-fused equivalent path
- shim transformers symbols moved to pytorch_utils
- MSVD t2v/v2t ndcg@10 0.859/0.858, mock-run 37/37
-
refactor: drop dead peft remap, fix install hint, add n_embedding_parameters
-
refactor: drop dead peft remap, fix install hint, add n_embedding_parameters
-
chore: declare internvideo2 as a conflicting extra, regenerate uv.lock
-
refactor: load InternVideo2-CLIP via SentenceTransformers
-
style: fix InternVideo2 lint
-
Update mteb/models/model_implementations/internvideo2_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
- chore: drop mock run results from the diff
Co-authored-by: Hubert Lu <hubielu@email.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (2902127)