Skip to content

2.12.27

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 23 Apr 21:58
· 739 commits to main since this release

2.12.27 (2026-04-23)

Fix

  • fix: normalize package name (#4491)

Apparently pipy normalizes package name which gives a probleM that only appear when install from pipy:

ValueError: Unknown extras group(s) for mteb: ['google_genai']. Available: ['ark', 'audio', 'blip2', 'bm25s', 'codecarbon', 'cohere', 'colpali-engine', 'colqwen3', 'eager-embed', 'embeddinggemma', 'faiss-cpu', 'flagembedding', 'flash-attention', 'google-genai', 'gritlm', 'image', 'jina', 'jina-clip', 'jina-v4', 'leaderboard', 'llama-embed-nemotron', 'llama-nemotron-colembed-vl', 'llama-nemotron-embed-vl-1b-v2', 'llm2vec', 'mctct', 'model2vec', 'msclap', 'muq', 'nemotron-colembed-vl-v2', 'nomic', 'open-clip-torch', 'openai', 'peft', 'pylate', 'qwen-omni-utils', 'qwen-vl', 'sauerkrautlm-colpali', 'siglip', 'speechbrain', 'timm', 'torch-vggish-yamnet', 'vertexai', 'video', 'vllm', 'voyage-v', 'voyageai', 'wav2clip', 'xet', 'xformers', 'youtu']

pep: https://peps.python.org/pep-0685/

initially proposed a fix here, but discovered that it was added:
#4384

in this commit:
45f1419

This PR just bumps the version to release the fix.

(cc @isaac-chung)

Co-authored-by: Kenneth <kennethenevoldsen@gmail.com> (f3f242d)

Unknown

  • add mteb/AudioCaps_AV dataset (#4478)

  • add mteb/AudioCaps_AV dataset

  • Add import for AVMemeExam retrieval classes

  • fix license, domains, date, and va2t/t2va prompts for AudioCaps-AV

  • license: mit -> cc-by-4.0 (AudioCaps is built on AudioSet, CC-BY 4.0)
  • domains: Encyclopaedic -> AudioScene (YouTube audio scene clips)
  • date: 2018 -> 2019 (NAACL 2019 paper)
  • va2t prompt: mention audio alongside video
  • t2va prompt: mention audio alongside video

Made-with: Cursor

  • fix AudioCaps-AV descriptions and prompts to reflect audio-only captions

AudioCaps_AV has a single caption field containing human-written audio scene
descriptions (what is heard, not seen). Descriptions now clarify the
cross-modal nature of v2t/t2v and that va2t/t2va are audio-focused.
Prompts updated to match: 'sounds in', 'audio description', 'what is heard'.

Made-with: Cursor

  • Revert "fix AudioCaps-AV descriptions and prompts to reflect audio-only captions"

This reverts commit 9139801.

  • Revert "fix license, domains, date, and va2t/t2va prompts for AudioCaps-AV"

This reverts commit 97a9e49.


Co-authored-by: AdnanElAssadi56 <115242814+AdnanElAssadi56@users.noreply.github.com> (0b71456)

  • [MVEB] Add Qwen Omni Video Support (#4412)

  • fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification

  • fix: Move SIBFLEURS descriptive stats to AudioClassification

  • feat: add video support to Qwen Omni models and remove qwen_omni_utils dependency

Add video modality to QwenOmniWrapper by handling video frames directly
via the processor instead of going through qwen_omni_utils.process_mm_info,
which only supports URL/path loading and crashes on pre-loaded tensors.
Video frames are pre-resized via smart_resize to match the model's expected
pixel range. Also fixes dtype mismatch for Qwen3 models and updates
max_audio_length to 300s to match the model's preprocessor config (n_samples).

  • feat: add video support to Qwen Omni models and remove qwen_omni_utils dependency

Add video modality to QwenOmniWrapper by handling video frames directly
via the processor instead of going through qwen_omni_utils.process_mm_info,
which only supports URL/path loading and crashes on pre-loaded tensors.
Video frames are pre-resized via smart_resize to match the model's expected
pixel range. Also fixes dtype mismatch for Qwen3 models and updates
max_audio_length to 300s to match the model's preprocessor config (n_samples).

  • refactor: pass video/audio directly to processor, use FPS sampling
  • Remove _resize_video (processor handles resize via its own config)
  • Remove manual audio extraction loop (collator handles resampling)
  • Pass tensors directly to processor with do_sample_frames=False
  • Use fps=2.0 default for video frame subsampling in collator
  • Use AudioCollator for audio-only tasks
  • fix: remove unused extra_requirements_groups for qwen_omni_utils

  • fix: remove unnecessary dtype cast per reviewer feedback (857b425)

  • add image to LCO embed model (#4384)

  • add image and video to LCO embed model

  • Fix extras group comparison to be underscore/hyphen insensitive

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>


Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> (fb64050)

  • add mteb/panda-70m dataset (#4469)

  • add-panda-70m-dataset

  • fix lint (24365e9)

  • add mteb/AVMeme-Exam dataset (#4480) (25991e6)