2.12.19
2.12.19 (2026-04-16)
Documentation
-
docs: Update adding dataset checklist (#4394)
-
docs: Update adding dataset checklist
fix the checklist to make it less text-specific
- docs: Add score reproduction to PR reqirenments (#4396)
add score reproduction to description
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (dc58c76)
Fix
-
fix: Auto add base model to
ModelMeta(#4395) -
fetch source model from hub
-
fix tests
-
check if model has model card attr (
ed1833a)
Unknown
-
model: Add Google Gemini embedding 2 (#4247)
-
Adding Google Gemini embedding 2 model
-
feat: add per-task prompt mapping and multimodal support for Gemini Embedding 2
- Create GEMINI_EMBEDDING_2_PROMPTS dict with 132 per-task Google API
task type mappings from issue #4260 - Add GoogleGeminiEmbeddingModel class using google-genai SDK with
support for text, image, and interleaved text+image inputs - Update ModelMeta to use new class, set modalities=["image", "text"]
- Add google_genai optional dependency to pyproject.toml
- feat: add audio modality support for Gemini Embedding 2
- Add _audio_to_wav_bytes helper to convert numpy audio arrays to WAV
- Handle audio inputs in encode() via Part.from_bytes with audio/wav MIME
- Update modalities to ["audio", "image", "text"]
- fix: strip google/ prefix from model name for Gemini API
The google-genai SDK's embed_content doesn't handle the "google/"
prefix format. Strip it in the constructor like Voyage does.
- fix: add exponential backoff retry for 429 rate limits
Retry up to 10 times with exponential backoff (60s, 120s, 240s...
up to 600s) when hitting API quota limits. Essential for large
multilingual benchmarks like MIRACL.
- refactor: address PR review comments
- Replace 132-entry per-task dict with task-type defaults + 62
per-task overrides (KennethEnevoldsen: use metadata) - Add embed_dim parameter to GoogleGeminiEmbeddingModel (Samoed)
- Add title formatting for retrieval corpus docs (Samoed)
- Add batch size comment referencing API limits (Samoed)
- Simplify encode() control flow
-
fix: replace print with logger.warning for lint compliance
-
fix: handle audio+text interleaved input and note MRL embed_dim support
- Add audio+text branch in encode() for interleaved content
- Note embed_dim supports [768, 1536, 3072] once PR #4170 is merged
-
fix: use MRL embed_dim list and remove duplicate logger
-
fix unused param
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (633e41c)
-
Move Kinetics400 out of video and add zeroshot version (#4383)
-
distinguish between AV and V tasks
-
move out of video folder and add zeroshot version
-
fix task type
-
update task metadata based on discussions
-
fix mveb task type mapping
-
fix: Add
is_betato task metadata (#4392) -
fix: Add
is_betato task metadata
- Added *is_beta to the task metadata
- Added a warning on initializing a dataset when is_beta is True
- Added exclude_beta to get_tasks and filter_tasks, for now I set it to False
todo:
- add tests
-
add test and updates metadata
-
format
-
re-enable tests for beta datasets
-
format
-
feat: comment out MVEB task types without existing tasks
VideoClustering, VideoPairClassification, and VideoCentricQA are defined
in task_metadata but have no corresponding task implementations yet,
causing create_available_tasks.py to fail. Comment them out until tasks
are added. Also regenerate available_tasks docs and add qwen_omni_utils
optional dependency.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- fix: allow beta-only task types in create_available_tasks.py
Change assertion to <= so task types that only have beta tasks don't
break the docs generation. Use .get() with continue to skip task types
with no non-beta tasks.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- fix: skip video tasks missing descriptive_stats in metadata test
Skip Kinetics400 video tasks in test_all_metadata_is_filled_and_valid
until descriptive stats are added. Regenerate available_tasks docs.
-
revert: restore docs/overview/available_tasks to main
-
revert: remove all generated available_tasks changes from branch
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> (6ec3f40)
-
Add nicher92/saga-embed_v1 to MTEB models (#4371)
-
Add nicher92/saga-embed_v1 to MTEB models
-
Update training_datasets in ModelMeta
-
fix: fixed naming
-
Replace custom SagaModel class with standard SentenceTransformerEncoderWrapper and model_prompts dict
-
chore: remove lingering comment
-
Update mteb/models/model_implementations/saga_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
-
update meta
-
change parameters and memory usage
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (d7c521c)