2.12.17
·
786 commits
to main
since this release
2.12.17 (2026-04-16)
Fix
- fix: Corrected incorrect model rename (#4391)
This gave the following incorrect warning:
DeprecationWarning: The model 'mteb/baseline-random-encoder' has been renamed to 'mteb/baseline-random-encoder'. To prevent this warning use the new name.
model = mteb.get_model_meta("mteb/baseline-random-encoder") ([`b65730d`](https://github.com/embeddings-benchmark/mteb/commit/b65730d833b3e321be759b243f62090066faad45))
## Unknown
* model: add BidirLM text embedding family (270M, 0.6B, 1B, 1.7B) (#4374)
* model: add BidirLM text embedding family (270M, 0.6B, 1B, 1.7B)
* Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
* run lint
---------
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> ([`e8a4069`](https://github.com/embeddings-benchmark/mteb/commit/e8a40693b00005cb0362bd6c5798d61287a196f3))
* [MVEB] PE-AV Model, Kinetics400 Dataset, RavdessAV Dataset (#4199)
* fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification
* fix: Move SIBFLEURS descriptive stats to AudioClassification
* Adding video modality
* Add Kinetics-400 dataset
* Add pe_av model
* fix typo
* fix collator bug
* Edit selecting column in classification abstask
* Properly handle frames in PE_AV
* add self kwarg to method
* Add audio collator
* fix type error
* fix audio_video embeds object handling
* Add Ravdess_av clustering
* fix task metadata
* start video integration
* start video integration
* upd task structure
* upd video input type
* combine video and audio to dict
* fix task side
* fix pe_av model
* lower writer batch size
* fix col labels
* lint
* add pe_av model metadata
* fix datasets metadata
* remove accidently commited files
* remove nested list structure from datasets
* edit collator to handle one video item
* multimodal collator + fix comments
* lint
* metadata update
* using forward pass to get embeds
* replace forward pass + add audio to msrvtt
* fix category metadata
* edit get embeddings
* add n_embedding_parameters
* change input col name to list
* lint + type check
* add classvar
* add str to classvar
* Change list to sequence
* lint + type check error
* edit dataloader and msrvtt handling of input column
* move seqeuence out of type checking
* fix random baseline
* add collator to random baseline
* restore previous dict structure + make audio optional
* clean structure
* lint
* safety check
* decrease writer batch size
* match msrvtt format
* type check fix
* refactor: keep video and audio as separate dataset columns
* fix: handle single-string input_column correctly in _prepare_dataset
* review fixes
* lint
* type hins fix
* address review: simplify input_column_name, remove VideoInputItem, fix collator output
- Revert input_column_name from Mapping[str, str] to str | Sequence[str]
- Remove VideoInputItem wrapper, pass frames tensor directly
- Make VideoCollator return BatchedInput (consistent with AudioCollator)
- MultimodalCollator uses static methods instead of chaining collators
* fix: update clustering_evaluator to use Sequence instead of Mapping
* fix: handle Sequence input_column_name in second create_dataloader call
* fix: skip statistics and text cleaning for multi-column video tasks
* fix: pass explicit None for TypedDict fields in multi-column statistics
* address Kenneth review: rename collators, update docs, simplify annotations
- Rename VideoCollator -> FramesCollator, MultimodalCollator -> VideoCollator
- Update VideoInput docstring to clarify frames-only, audio in AudioInput
- Update input_column_name docs in classification/clustering base classes
- Use ClassVar[Sequence[str]] for video task input_column_name
- Extract isinstance check to top of zeroshot evaluator __call__
- Improve task_pipelines.py skip comment for multi-column tasks
- Add TODO for MSR-VTT dataset reupload
* docs: link to encoder I/O types for default column names in input_column_name
* fix: raise NotImplementedError for multi-column task cleaning
* refactor: use tuples for input_column_name to avoid ClassVar
* refactor: move Sequence handling into create_dataloader, simplify callers
---------
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> ([`5d3c845`](https://github.com/embeddings-benchmark/mteb/commit/5d3c8453db615a1a4ae4d53033711af21c1d502e))
* dataset: add BrowseComp-Plus (#4226)
* dataset: Add BrowseComp-Plus
* fix linting errors
* fixing bibtext formatting
* Split BrowseCompPlusRetrieval into gold_only and gold_and_evidence subsets
* fix: remove qa as a valid tag for metadata files
* simplify data loading by reuploading the data
---------
Co-authored-by: Kenneth <kennethenevoldsen@gmail.com> ([`e722b76`](https://github.com/embeddings-benchmark/mteb/commit/e722b7640ed1abee68c3df5023a186b36a15325f))