Repository navigation
2.12.27
2.12.27 (2026-04-23)
Fix
- fix: normalize package name (#4491)
Apparently pipy normalizes package name which gives a probleM that only appear when install from pipy:
ValueError: Unknown extras group(s) for mteb: ['google_genai']. Available: ['ark', 'audio', 'blip2', 'bm25s', 'codecarbon', 'cohere', 'colpali-engine', 'colqwen3', 'eager-embed', 'embeddinggemma', 'faiss-cpu', 'flagembedding', 'flash-attention', 'google-genai', 'gritlm', 'image', 'jina', 'jina-clip', 'jina-v4', 'leaderboard', 'llama-embed-nemotron', 'llama-nemotron-colembed-vl', 'llama-nemotron-embed-vl-1b-v2', 'llm2vec', 'mctct', 'model2vec', 'msclap', 'muq', 'nemotron-colembed-vl-v2', 'nomic', 'open-clip-torch', 'openai', 'peft', 'pylate', 'qwen-omni-utils', 'qwen-vl', 'sauerkrautlm-colpali', 'siglip', 'speechbrain', 'timm', 'torch-vggish-yamnet', 'vertexai', 'video', 'vllm', 'voyage-v', 'voyageai', 'wav2clip', 'xet', 'xformers', 'youtu']
pep: https://peps.python.org/pep-0685/
initially proposed a fix here, but discovered that it was added:
#4384
in this commit:
45f1419
This PR just bumps the version to release the fix.
(cc @isaac-chung)
Co-authored-by: Kenneth <kennethenevoldsen@gmail.com> (f3f242d)
Unknown
-
add mteb/AudioCaps_AV dataset (#4478)
-
add mteb/AudioCaps_AV dataset
-
Add import for AVMemeExam retrieval classes
-
fix license, domains, date, and va2t/t2va prompts for AudioCaps-AV
- license: mit -> cc-by-4.0 (AudioCaps is built on AudioSet, CC-BY 4.0)
- domains: Encyclopaedic -> AudioScene (YouTube audio scene clips)
- date: 2018 -> 2019 (NAACL 2019 paper)
- va2t prompt: mention audio alongside video
- t2va prompt: mention audio alongside video
Made-with: Cursor
- fix AudioCaps-AV descriptions and prompts to reflect audio-only captions
AudioCaps_AV has a single caption field containing human-written audio scene
descriptions (what is heard, not seen). Descriptions now clarify the
cross-modal nature of v2t/t2v and that va2t/t2va are audio-focused.
Prompts updated to match: 'sounds in', 'audio description', 'what is heard'.
Made-with: Cursor
- Revert "fix AudioCaps-AV descriptions and prompts to reflect audio-only captions"
This reverts commit 9139801.
- Revert "fix license, domains, date, and va2t/t2va prompts for AudioCaps-AV"
This reverts commit 97a9e49.
Co-authored-by: AdnanElAssadi56 <115242814+AdnanElAssadi56@users.noreply.github.com> (0b71456)
-
[MVEB] Add Qwen Omni Video Support (#4412)
-
fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification
-
fix: Move SIBFLEURS descriptive stats to AudioClassification
-
feat: add video support to Qwen Omni models and remove qwen_omni_utils dependency
Add video modality to QwenOmniWrapper by handling video frames directly
via the processor instead of going through qwen_omni_utils.process_mm_info,
which only supports URL/path loading and crashes on pre-loaded tensors.
Video frames are pre-resized via smart_resize to match the model's expected
pixel range. Also fixes dtype mismatch for Qwen3 models and updates
max_audio_length to 300s to match the model's preprocessor config (n_samples).
- feat: add video support to Qwen Omni models and remove qwen_omni_utils dependency
Add video modality to QwenOmniWrapper by handling video frames directly
via the processor instead of going through qwen_omni_utils.process_mm_info,
which only supports URL/path loading and crashes on pre-loaded tensors.
Video frames are pre-resized via smart_resize to match the model's expected
pixel range. Also fixes dtype mismatch for Qwen3 models and updates
max_audio_length to 300s to match the model's preprocessor config (n_samples).
- refactor: pass video/audio directly to processor, use FPS sampling
- Remove _resize_video (processor handles resize via its own config)
- Remove manual audio extraction loop (collator handles resampling)
- Pass tensors directly to processor with do_sample_frames=False
- Use fps=2.0 default for video frame subsampling in collator
- Use AudioCollator for audio-only tasks
-
fix: remove unused extra_requirements_groups for qwen_omni_utils
-
fix: remove unnecessary dtype cast per reviewer feedback (
857b425) -
add image to LCO embed model (#4384)
-
add image and video to LCO embed model
-
Fix extras group comparison to be underscore/hyphen insensitive
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> (fb64050)