Repository navigation
2.12.28
2.12.28 (2026-04-25)
Fix
-
fix: Added Git Actions using Command Pattern (#4329)
-
Added Git Actions using Command Pattern
-
lintter
-
Added better error logging in case of rollback
-
added init.py and github group
-
update docstring of function based on copilot comment
-
update docstring of PushToFork
-
fix dependencies
-
add in Makefile install and fix import
-
fix GithubException import
-
fix install command in Makefile
-
apply changes from review
-
update description of CopyResultsAction
-
added github in install-for-tests
-
make module private
-
make folder private
-
Fix undo in case of failure
-
remove CopyResultsAction
-
fix imports
-
added init.py
-
fix import
-
fix lintter
-
remove comments
-
Added pytest.importorskip
-
fix lint
-
fix lint
-
Remove monkeypatch from tests and update tests
-
fix default branch in test
-
fix test and cleanup
Co-authored-by: Copilot <copilot@github.com>
-
apply changes from review
-
setting email and username in config only when not set
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Copilot <copilot@github.com> (b2dfda9)
Refactor
- refactor: replace process_mm_info with direct processor calls for LCO
Add VideoCollator/AudioCollator for proper video/audio support,
remove qwen_omni_utils dependency, add L2 normalization, and
expose fps/max_frames/num_frames/max_audio_length params. (eae2080)
Style
- style: lint formatting (
b804411)
Unknown
- fix ci after video statistics merge (#4510)
fix ci (baa25a9)
-
[MVEB] Add video statistics (#4456)
-
start video statistics
-
add some more video processing
-
more changes
-
simplify statistics calculation
-
refactor updates
-
update mock video tasks
-
fix typing
-
make all tasks support multimodality
-
add minimal version for av
-
fix sts
-
fix tests
-
update skip rule
-
activate video statistics check
-
add missing tasks
-
move statistics functions
-
update column format
-
upd lock
-
add resolutions
-
add to ignore newer datasets (
6d862aa) -
add remaining v2a and a2v tasks for video retrieval datasets (#4504) (
0294578) -
add v2a and a2v tasks for valid video retrieval datasets (#4494)
-
add v2a and a2v tasks to MSR-VTT
Made-with: Cursor
- add v2a and a2v tasks to DiDeMo
Made-with: Cursor
- add v2a and a2v tasks to YouCook2
Made-with: Cursor
- add v2a and a2v tasks to Shot2Story20K
Made-with: Cursor
- add v2a and a2v tasks to VALOR-32K
Made-with: Cursor
- add v2a and a2v tasks to VATEX
Made-with: Cursor
- add 'Cross-Modal Retrieval' to TaskSubtype allowed values
Required for v2a and a2v retrieval tasks across video datasets
(DiDeMo, YouCook2, Shot2Story20K, VALOR-32K, VATEX, MSR-VTT).
Made-with: Cursor (6119afb)
-
Fix an incorrect retrieval example in docs (#4496)
-
Fix an incorrect retrieval example in docs
Incorrect assignment of the data to self.dataset. It is None, should be reinitialized
- format
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (cf4a7bf)
- Merge branch 'mveb-lco-video-support'
Conflicts:
mteb/models/model_implementations/lco_embedding_models.py (f46321b)
-
[MVEB] Move collators to separate file (#4482)
-
move collators
-
make parameters by name
-
fix last import (
25a0991) -
[MVEB] Omni Embed Nemotron (#4388)
-
fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification
-
fix: Move SIBFLEURS descriptive stats to AudioClassification
-
add: NVIDIA Omni-Embed-Nemotron-3B model wrapper
-
fix: convert video frame tensors to PIL images for qwen_omni_utils
-
fix: add default system prompt to suppress Qwen audio output warning
-
fix: correct text max_length to 204800 per model card
-
refactor: remove qwen_omni_utils dependency from omni-embed-nemotron wrapper
- Remove process_mm_info and pass media directly to processor
- Remove PIL conversion, pass video tensors directly to processor
- Reimplement smart_resize pre-step to match model card behavior
- Add images_kwargs per paper Table 7 (min_pixels=196, max_pixels=2352)
- Fix duplicate padding kwarg (transformers 5.x compat)
-
fix: address PR review - lint, vidore training datasets, lowercase vars
-
refactor: use SentenceTransformerMultimodalEncoderWrapper instead of custom wrapper
-
refactor: add processing_kwargs for video/audio in ST wrapper
-
fix: support ST multimodal wrapper for omni-embed-nemotron
- _prepare_dataset drops non-modality columns (fixes ST 5.4 strict validation)
- SentenceTransformerMultimodalEncoderWrapper uses np.concatenate to flatten batch dim
- add omni-embed-nemotron extra for sentence-transformers>=5.4 / transformers>=5.3
- add: RAVDESS + MSRVTT results for omni-embed-nemotron (ST wrapper)
RAVDESSAVClustering v_measure: 0.0794
MSRVTTV2T ndcg_at_10: 0.3196
-
change placement
-
refactor: filter non-modality keys in ST multimodal wrapper instead of per-task
-
refactor: move audio unwrap to ST base wrapper, simplify omni-embed-nemotron
- Move AudioInputItem -> raw array unwrap from the model-specific
encode() into SentenceTransformerMultimodalEncoderWrapper.encode()
so all multimodal ST models benefit (guarded byif "audio" in batch).
Works around sentence-transformers#3732. - Drop num_frames parameter from OmniEmbedNemotronWrapper and rely on
VideoCollator defaults. - Rename pyproject extra omni-embed-nemotron -> multimodal_sbert since
the deps (sentence-transformers>=5.4.0, transformers>=5.3.0) are
generic to multimodal ST models, not specific to this one. - Re-run RAVDESS + MSRVTT with the refactored wrapper to confirm parity
(MSRVTT stable, RAVDESS drifted within BF16-noise envelope).
- add: extra_requirements_groups for omni-embed-nemotron
With the merge of upstream/main (PR #4356), ModelMeta now supports
extra_requirements_groups. Declare multimodal_sbert so the runtime
enforces the pyproject pins at load time.
- refactor: move collator params to SentenceTransformerMultimodalEncoderWrapper
- Add fps, max_frames, num_frames, target_sampling_rate, max_samples
to the multimodal wrapper init so subclasses don't need to override
encode() just to set a collator - Auto-set VideoCollator for video tasks, AudioCollator for audio-only
- Simplify OmniEmbedNemotronWrapper: remove encode() override, pass
collator params via super().init() instead - Set fps=2.0 for omni-embed-nemotron to avoid decoding all frames
(do_sample_frames=False means processor won't resample)
-
fix: simplify multimodal wrapper docstring
-
fix tests
-
partly fix typing
-
add: num_frames param to OmniEmbedNemotronWrapper
-
add: e5-omni-3B and e5-omni-7B model implementations
-
fix: use hyphenated extra name multimodal-sbert for pip compatibility
-
fix: normalize multimodal_sbert to multimodal-sbert in pyproject.toml
-
remove: e5-omni models and results per reviewer request
Moving e5-omni to a separate PR to not block this one.
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4a319fa)
-
add mteb/VGGSound_AV_RETRIEVAL dataset (#4479)
-
add mteb/VGGSound_AV_RETRIEVAL dataset
-
update to video captions
-
refine VGGSound-AV: guard av_caption map and improve descriptions/prompts
- Only compute av_caption merge when the task actually needs it (va2t/t2va),
avoiding a full dataset.map pass for v2t and t2v tasks - va2t/t2va descriptions now use natural language covering both modalities
- Prompts for va2t/t2va explicitly mention what is seen and heard
Made-with: Cursor (29c2a27)