Skip to content

2.18.7

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 26 Jul 12:38
· 360 commits to main since this release

2.18.7 (2026-07-26)

Ci

  • ci: update ruff to 0.16 (#5008)

update ruff to 0.16 (df38f9b)

Fix

  • fix: bm25s fall back to language-agnostic tokenization for unknown langs (#5009)

fix(bm25): fall back to language-agnostic tokenization for unknown languages

Co-authored-by: Nicolas Helmeyer <helmeyen@login-4.server.mila.quebec> (787bf7c)

Unknown

  • Leaderboard: Add MTEB(por, v1) to homepage language specific. (#4970)

add mteb-pt to homepage (8847790)

  • model: add webAI ColVec1.1 4B and 8B models (#5010)

  • model: add webAI ColVec1.1 models

  • model: declare ColVec1.1 transformers requirement

  • fix(model): address initial ColVec1.1 review feedback

  • fix(model): make SDPA the ColVec1.1 default (6e72309)

  • model: Update fusion-embedding-2 revision to v0.3-preview (1720d8b1) (#5028) (46a2c21)

  • Add ModelMeta: minetta/nemotron-3-embed-8b-legal (#5027)

  • Add ModelMeta: minetta/nemotron-3-embed-8b-legal

  • license as URL (openmdw-1.1 not in Licenses literal)

  • Fill n_embedding_parameters; set public_training_code/data explicitly

  • revert public_training_* to None (schema expects str|None)


Co-authored-by: banyaneth <banyaneth@users.noreply.github.com> (6089992)

  • Add KiteFishAI/Nano-Em1-0.6B-v2 (#4993)

  • Add KiteFishAI/Nano-Em1-0.6B-v2

  • Update model revision hash in qwen3_models.py

  • Update Nano-Em1-0.6B-v2 model metadata

Updated model metadata for Nano-Em1-0.6B-v2 including release date, parameter counts, and memory usage.

  • Update qwen3_models.py

  • Refactor training datasets format in qwen3_models.py

  • Update qwen3_models.py

  • Add ScoringFunction import to qwen3_models.py

  • format & move

  • format & move


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (c67d000)

  • Add UrbanSound8K Audio Clustering task (closes #5018) (#5025)

Part of MOEB: Massive Omni Embedding Benchmark (tracking issue #4842)

Co-authored-by: Rakshitha Ireddi <rakshithaireddi@Rakshithas-MacBook-Pro.local> (609b43f)

  • task: add Song Describer text-music retrieval (t2a, a2t) (#4988)

  • task: add Song Describer text-music retrieval (t2a, a2t)

Human-written music captions from the Song Describer Dataset (CC-BY-SA-4.0,
MTG-Jamendo audio) as text<->music retrieval. 746 captions over 547 tracks.
CLAP t2a hit_rate@5 6.3 vs random 1.3. Complements MusicCaps with human captions
and a published retrieval benchmark. Includes descriptive statistics.

  • task: use full Song Describer release (706 tracks / 1106 captions)

Rebuild from the full Zenodo SDD release (was the 547-track valid subset), matching
the paper's corpus. laion/larger_clap_general reproduces the paper's Table 5 CLAP
retrieval almost exactly (T2A R@1/5/10 = 4.79/17.72/29.48 vs paper 4.42/17.02/26.01).
Updated revisions, descriptions, and descriptive statistics.

  • Update mteb/tasks/retrieval/zxx/song_describer.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • task: also use recall_at_5 for Song Describer A2T (consistency)

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (d206dae)

  • Update links to code for Nemotron models (#5002)

update links to code for Nemotron models (18005a3)

  • [MOEB] Add AESDD dataset (#4978)

  • [MOEB] Add AESDD dataset

  • remove corrupt audiofile

  • simplifying implementation with AESDD fixed (8751596)

  • task: add CASTELLA audio moment retrieval (t2a) (#4984)

  • task: add CASTELLA audio moment retrieval (t2a)

First temporal-localization retrieval task for the audio benchmark:
captions retrieve the 10-second window of a long recording containing
the described moment, built on CASTELLA (arXiv 2511.15131), the DCASE
2026 Task 6 evaluation set. 12,046 windows from 566 recordings, 1,347
queries, graded by >=50 percent overlap with annotated moments.

  • chore: add descriptive statistics for CASTELLAAMRRetrieval

  • task: switch CASTELLA-AMR to full-recording retrieval

Replace the 10s-window corpus with the 566 complete recordings (60-300s),
one gold recording per caption. Matches the paper's audio length and simplifies
the task; recomputed descriptive statistics.

  • style: ruff format castella_amr (885f740)

  • model: add SigLIP2 family (15 checkpoints) (#4973)

  • model: add SigLIP2 family (15 checkpoints)

Adds ModelMeta entries for the google/siglip2-* fixed-resolution
checkpoints (base/large/so400m/giant-opt). They reuse the existing
SiglipModelWrapper since these checkpoints load as SiglipModel.
NaFlex variants are excluded as they need different processor handling.

Closes #2301

  • review: set languages and drop unneeded extras for SigLIP2

SigLIP2 ships a fast tokenizer so sentencepiece/protobuf are not
required; the image extra is added automatically. (62c5293)

  • [MOEB]: Add Covers80 dataset (#4986)

  • [MOEB]: Add Covers80 dataset

  • update description

  • formatting

  • add descriptive stats

  • small improvement script (be0f623)

  • Add Hanno-Labs/dinghy-law-4b-v1 (legal embedding model) (#4992)

  • Add Hanno-Labs/dinghy-law-4b-v1 (legal embedding model)

  • Apply suggestions from code review

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (41a5310)