2.18.7
2.18.7 (2026-07-26)
Ci
- ci: update ruff to 0.16 (#5008)
update ruff to 0.16 (df38f9b)
Fix
- fix: bm25s fall back to language-agnostic tokenization for unknown langs (#5009)
fix(bm25): fall back to language-agnostic tokenization for unknown languages
Co-authored-by: Nicolas Helmeyer <helmeyen@login-4.server.mila.quebec> (787bf7c)
Unknown
- Leaderboard: Add MTEB(por, v1) to homepage language specific. (#4970)
add mteb-pt to homepage (8847790)
-
model: add webAI ColVec1.1 4B and 8B models (#5010)
-
model: add webAI ColVec1.1 models
-
model: declare ColVec1.1 transformers requirement
-
fix(model): address initial ColVec1.1 review feedback
-
fix(model): make SDPA the ColVec1.1 default (
6e72309) -
model: Update fusion-embedding-2 revision to v0.3-preview (1720d8b1) (#5028) (
46a2c21) -
Add ModelMeta: minetta/nemotron-3-embed-8b-legal (#5027)
-
Add ModelMeta: minetta/nemotron-3-embed-8b-legal
-
license as URL (openmdw-1.1 not in Licenses literal)
-
Fill n_embedding_parameters; set public_training_code/data explicitly
-
revert public_training_* to None (schema expects str|None)
Co-authored-by: banyaneth <banyaneth@users.noreply.github.com> (6089992)
-
Add KiteFishAI/Nano-Em1-0.6B-v2 (#4993)
-
Add KiteFishAI/Nano-Em1-0.6B-v2
-
Update model revision hash in qwen3_models.py
-
Update Nano-Em1-0.6B-v2 model metadata
Updated model metadata for Nano-Em1-0.6B-v2 including release date, parameter counts, and memory usage.
-
Update qwen3_models.py
-
Refactor training datasets format in qwen3_models.py
-
Update qwen3_models.py
-
Add ScoringFunction import to qwen3_models.py
-
format & move
-
format & move
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (c67d000)
Part of MOEB: Massive Omni Embedding Benchmark (tracking issue #4842)
Co-authored-by: Rakshitha Ireddi <rakshithaireddi@Rakshithas-MacBook-Pro.local> (609b43f)
-
task: add Song Describer text-music retrieval (t2a, a2t) (#4988)
-
task: add Song Describer text-music retrieval (t2a, a2t)
Human-written music captions from the Song Describer Dataset (CC-BY-SA-4.0,
MTG-Jamendo audio) as text<->music retrieval. 746 captions over 547 tracks.
CLAP t2a hit_rate@5 6.3 vs random 1.3. Complements MusicCaps with human captions
and a published retrieval benchmark. Includes descriptive statistics.
- task: use full Song Describer release (706 tracks / 1106 captions)
Rebuild from the full Zenodo SDD release (was the 547-track valid subset), matching
the paper's corpus. laion/larger_clap_general reproduces the paper's Table 5 CLAP
retrieval almost exactly (T2A R@1/5/10 = 4.79/17.72/29.48 vs paper 4.42/17.02/26.01).
Updated revisions, descriptions, and descriptive statistics.
- Update mteb/tasks/retrieval/zxx/song_describer.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
- task: also use recall_at_5 for Song Describer A2T (consistency)
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (d206dae)
- Update links to code for Nemotron models (#5002)
update links to code for Nemotron models (18005a3)
-
[MOEB] Add AESDD dataset (#4978)
-
[MOEB] Add AESDD dataset
-
remove corrupt audiofile
-
simplifying implementation with AESDD fixed (
8751596) -
task: add CASTELLA audio moment retrieval (t2a) (#4984)
-
task: add CASTELLA audio moment retrieval (t2a)
First temporal-localization retrieval task for the audio benchmark:
captions retrieve the 10-second window of a long recording containing
the described moment, built on CASTELLA (arXiv 2511.15131), the DCASE
2026 Task 6 evaluation set. 12,046 windows from 566 recordings, 1,347
queries, graded by >=50 percent overlap with annotated moments.
-
chore: add descriptive statistics for CASTELLAAMRRetrieval
-
task: switch CASTELLA-AMR to full-recording retrieval
Replace the 10s-window corpus with the 566 complete recordings (60-300s),
one gold recording per caption. Matches the paper's audio length and simplifies
the task; recomputed descriptive statistics.
-
style: ruff format castella_amr (
885f740) -
model: add SigLIP2 family (15 checkpoints) (#4973)
-
model: add SigLIP2 family (15 checkpoints)
Adds ModelMeta entries for the google/siglip2-* fixed-resolution
checkpoints (base/large/so400m/giant-opt). They reuse the existing
SiglipModelWrapper since these checkpoints load as SiglipModel.
NaFlex variants are excluded as they need different processor handling.
Closes #2301
- review: set languages and drop unneeded extras for SigLIP2
SigLIP2 ships a fast tokenizer so sentencepiece/protobuf are not
required; the image extra is added automatically. (62c5293)
-
[MOEB]: Add Covers80 dataset (#4986)
-
[MOEB]: Add Covers80 dataset
-
update description
-
formatting
-
add descriptive stats
-
small improvement script (
be0f623) -
Add Hanno-Labs/dinghy-law-4b-v1 (legal embedding model) (#4992)
-
Add Hanno-Labs/dinghy-law-4b-v1 (legal embedding model)
-
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (41a5310)