Skip to content

2.12.28

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 25 Apr 11:47
· 725 commits to main since this release

2.12.28 (2026-04-25)

Fix

  • fix: Added Git Actions using Command Pattern (#4329)

  • Added Git Actions using Command Pattern

  • lintter

  • Added better error logging in case of rollback

  • added init.py and github group

  • update docstring of function based on copilot comment

  • update docstring of PushToFork

  • fix dependencies

  • add in Makefile install and fix import

  • fix GithubException import

  • fix install command in Makefile

  • apply changes from review

  • update description of CopyResultsAction

  • added github in install-for-tests

  • make module private

  • make folder private

  • Fix undo in case of failure

  • remove CopyResultsAction

  • fix imports

  • added init.py

  • fix import

  • fix lintter

  • remove comments

  • Added pytest.importorskip

  • fix lint

  • fix lint

  • Remove monkeypatch from tests and update tests

  • fix default branch in test

  • fix test and cleanup

Co-authored-by: Copilot <copilot@github.com>

  • apply changes from review

  • setting email and username in config only when not set


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Copilot <copilot@github.com> (b2dfda9)

Refactor

  • refactor: replace process_mm_info with direct processor calls for LCO

Add VideoCollator/AudioCollator for proper video/audio support,
remove qwen_omni_utils dependency, add L2 normalization, and
expose fps/max_frames/num_frames/max_audio_length params. (eae2080)

Style

Unknown

  • fix ci after video statistics merge (#4510)

fix ci (baa25a9)

  • [MVEB] Add video statistics (#4456)

  • start video statistics

  • add some more video processing

  • more changes

  • simplify statistics calculation

  • refactor updates

  • update mock video tasks

  • fix typing

  • make all tasks support multimodality

  • add minimal version for av

  • fix sts

  • fix tests

  • update skip rule

  • activate video statistics check

  • add missing tasks

  • move statistics functions

  • update column format

  • upd lock

  • add resolutions

  • add to ignore newer datasets (6d862aa)

  • add remaining v2a and a2v tasks for video retrieval datasets (#4504) (0294578)

  • add v2a and a2v tasks for valid video retrieval datasets (#4494)

  • add v2a and a2v tasks to MSR-VTT

Made-with: Cursor

  • add v2a and a2v tasks to DiDeMo

Made-with: Cursor

  • add v2a and a2v tasks to YouCook2

Made-with: Cursor

  • add v2a and a2v tasks to Shot2Story20K

Made-with: Cursor

  • add v2a and a2v tasks to VALOR-32K

Made-with: Cursor

  • add v2a and a2v tasks to VATEX

Made-with: Cursor

  • add 'Cross-Modal Retrieval' to TaskSubtype allowed values

Required for v2a and a2v retrieval tasks across video datasets
(DiDeMo, YouCook2, Shot2Story20K, VALOR-32K, VATEX, MSR-VTT).

Made-with: Cursor (6119afb)

  • Fix an incorrect retrieval example in docs (#4496)

  • Fix an incorrect retrieval example in docs

Incorrect assignment of the data to self.dataset. It is None, should be reinitialized

  • format

Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (cf4a7bf)

  • Merge branch 'mveb-lco-video-support'

Conflicts:

mteb/models/model_implementations/lco_embedding_models.py (f46321b)

  • [MVEB] Move collators to separate file (#4482)

  • move collators

  • make parameters by name

  • fix last import (25a0991)

  • [MVEB] Adding HMDB51 Task (Clustering) (#4488) (8d381dc)

  • [MVEB] Omni Embed Nemotron (#4388)

  • fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification

  • fix: Move SIBFLEURS descriptive stats to AudioClassification

  • add: NVIDIA Omni-Embed-Nemotron-3B model wrapper

  • fix: convert video frame tensors to PIL images for qwen_omni_utils

  • fix: add default system prompt to suppress Qwen audio output warning

  • fix: correct text max_length to 204800 per model card

  • refactor: remove qwen_omni_utils dependency from omni-embed-nemotron wrapper

  • Remove process_mm_info and pass media directly to processor
  • Remove PIL conversion, pass video tensors directly to processor
  • Reimplement smart_resize pre-step to match model card behavior
  • Add images_kwargs per paper Table 7 (min_pixels=196, max_pixels=2352)
  • Fix duplicate padding kwarg (transformers 5.x compat)
  • fix: address PR review - lint, vidore training datasets, lowercase vars

  • refactor: use SentenceTransformerMultimodalEncoderWrapper instead of custom wrapper

  • refactor: add processing_kwargs for video/audio in ST wrapper

  • fix: support ST multimodal wrapper for omni-embed-nemotron

  • _prepare_dataset drops non-modality columns (fixes ST 5.4 strict validation)
  • SentenceTransformerMultimodalEncoderWrapper uses np.concatenate to flatten batch dim
  • add omni-embed-nemotron extra for sentence-transformers>=5.4 / transformers>=5.3
  • add: RAVDESS + MSRVTT results for omni-embed-nemotron (ST wrapper)

RAVDESSAVClustering v_measure: 0.0794
MSRVTTV2T ndcg_at_10: 0.3196

  • change placement

  • refactor: filter non-modality keys in ST multimodal wrapper instead of per-task

  • refactor: move audio unwrap to ST base wrapper, simplify omni-embed-nemotron

  • Move AudioInputItem -> raw array unwrap from the model-specific
    encode() into SentenceTransformerMultimodalEncoderWrapper.encode()
    so all multimodal ST models benefit (guarded by if &#34;audio&#34; in batch).
    Works around sentence-transformers#3732.
  • Drop num_frames parameter from OmniEmbedNemotronWrapper and rely on
    VideoCollator defaults.
  • Rename pyproject extra omni-embed-nemotron -> multimodal_sbert since
    the deps (sentence-transformers>=5.4.0, transformers>=5.3.0) are
    generic to multimodal ST models, not specific to this one.
  • Re-run RAVDESS + MSRVTT with the refactored wrapper to confirm parity
    (MSRVTT stable, RAVDESS drifted within BF16-noise envelope).
  • add: extra_requirements_groups for omni-embed-nemotron

With the merge of upstream/main (PR #4356), ModelMeta now supports
extra_requirements_groups. Declare multimodal_sbert so the runtime
enforces the pyproject pins at load time.

  • refactor: move collator params to SentenceTransformerMultimodalEncoderWrapper
  • Add fps, max_frames, num_frames, target_sampling_rate, max_samples
    to the multimodal wrapper init so subclasses don't need to override
    encode() just to set a collator
  • Auto-set VideoCollator for video tasks, AudioCollator for audio-only
  • Simplify OmniEmbedNemotronWrapper: remove encode() override, pass
    collator params via super().init() instead
  • Set fps=2.0 for omni-embed-nemotron to avoid decoding all frames
    (do_sample_frames=False means processor won't resample)
  • fix: simplify multimodal wrapper docstring

  • fix tests

  • partly fix typing

  • add: num_frames param to OmniEmbedNemotronWrapper

  • add: e5-omni-3B and e5-omni-7B model implementations

  • fix: use hyphenated extra name multimodal-sbert for pip compatibility

  • fix: normalize multimodal_sbert to multimodal-sbert in pyproject.toml

  • remove: e5-omni models and results per reviewer request

Moving e5-omni to a separate PR to not block this one.


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4a319fa)

  • add mteb/VGGSound_AV_RETRIEVAL dataset (#4479)

  • add mteb/VGGSound_AV_RETRIEVAL dataset

  • update to video captions

  • refine VGGSound-AV: guard av_caption map and improve descriptions/prompts

  • Only compute av_caption merge when the task actually needs it (va2t/t2va),
    avoiding a full dataset.map pass for v2t and t2v tasks
  • va2t/t2va descriptions now use natural language covering both modalities
  • Prompts for va2t/t2va explicitly mention what is seen and heard

Made-with: Cursor (29c2a27)