2.18.8
2.18.8 (2026-07-30)
Fix
Unknown
-
[MVEB] Add STARBench video-centric QA task (#4746)
-
[MVEB] Add STARBench video-centric QA task
-
[MVEB] Split STARBench into per-config tasks (feasibility, interaction, prediction, sequence)
-
fix: run ruff format on star_bench.py
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- task: add descriptive stats for STARBench classes
Co-authored-by: Yashwanth Devavarapu <yashwanthdevavarapu@Yashwanths-MacBook-Pro.local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> (a1e96c8)
-
model: add RTriever-4B and DIVER-Retriever-4B (#5039)
-
model: add RTriever-4B and DIVER-Retriever-4B
Two reasoning-intensive retrievers, both Qwen3-Embedding-4B fine-tunes, so both
reuse q3e_instruct_loader: last-token pooling, cosine similarity, and an
"Instruct: ...\nQuery:" prefix on queries only. That matches the template in
each checkpoint's config_sentence_transformers.json verbatim.
DIVER-Retriever-4B declares the BRIGHT corpora in training_datasets. Its model
card lists reasonir/reasonir-data among its training sets, and the hard-query
half of that dataset stores positive documents by id and reconstructs them from
xlangai/BRIGHT, so those corpora are part of its training data transitively.
RTriever-4B is trained on synthetic data with no MTEB task overlap.
- model: add the remaining DIVER-Retriever checkpoints
Adds Diver-Retriever-4B, -1.7B and -0.6B alongside the -1020 release. All four
share the same interface and reuse q3e_instruct_loader; the shared
training-data annotation is factored into DIVER_TRAINING_DATA.
The 1.7B checkpoint is built on the Qwen3-1.7B base LM rather than a
Qwen3-Embedding checkpoint. Its 1_Pooling/config.json carries the 4B model's
word_embedding_dimension (2560) while the checkpoint's hidden size is 2048, so
ModelMeta records the real value.
- model: declare the loader explicitly instead of reusing q3e_instruct_loader
Per review: borrowing another family's loader function couples these models to
the Qwen3 module, so a change there would silently reach RTriever and DIVER.
Each file now carries its own instruction_template and passes
InstructSentenceTransformerModel with explicit loader_kwargs.
The template is byte-identical to what q3e_instruct_loader produced
("Instruct: {instruction}\nQuery:" on queries, empty on documents), so scores
are unchanged. (604a73c)
- Updating GZTAN reference (#5047)
update GZTAN reference (d3f982e)
-
Add BirdCLEF Species Audio Clustering task (closes #5017) (#5023)
-
Add BirdCLEF Species Audio Clustering task (closes #5017)
Part of MOEB: Massive Omni Embedding Benchmark (tracking issue #4842)
-
fix: update BirdCLEF bibtex citation to Cañas et al. 2025
-
fix: correct bibtex field order for BirdCLEF citation
Co-authored-by: Rakshitha Ireddi <rakshithaireddi@Rakshithas-MacBook-Pro.local> (23b79d1)
-
[MOEB] Add MMVU dataset (#5038)
-
[MOEB] Add MMVU dataset
-
fix lowest dep issue and updating reference
-
simplify mmvu with push_to_hub, adding create data for reference
-
Apply suggestion from @Samoed
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (b7e6584)
-
[MOEB] Add SEA-VL: Multicultural VL Dataset for Southeast Asia (#5040)
-
[MOEB] Add SEA-VL: Multicultural VL Dataset for Southeast Asia
-
fix task and add stats
-
remove shs import
-
simplify sea_vl by pre-computing dataset (
9c30471) -
model: Change attention for codefuse models (#5032)
change attention for codefuse (2c772e6)
- Add ModelMeta for mixedbread-ai/deepset-mxbai-embed-de-large-v1 (#5044)
Registers the German-focused mxbai embedding model
(mixedbread-ai/deepset-mxbai-embed-de-large-v1, 487M params,
XLM-RoBERTa backbone, 1024-dim) in the model implementations
registry using the generic SentenceTransformerEncoderWrapper
reference loader.
This model already has 3.3M downloads on the Hub but was missing
from the MTEB registry, which blocks the pending results PR
(embeddings-benchmark/results#647) that reports its scores on the
full MTEB(deu, v1) benchmark.
Co-authored-by: Simon Dittrich <ai@cloudsoziologe.de> (a4046b9)
-
model: Add jina-reranker-v3.5 metadata (#5041)
-
model: add jina-reranker-v3.5 metadata
-
fix: support jina reranker v3.5 API
-
remove tests
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (e123184)
-
Dump vLLM to v0.26.0 (#5022)
-
init
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
-
refine
-
refine
-
- benchmarks and images
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
-
- benchmarks and images
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io> (fbdd40e)