Skip to content

2.11.0

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 21 Mar 19:20
· 879 commits to main since this release

2.11.0 (2026-03-21)

Feature

  • feat: Add Matryoshka support when loading a model (#4170)

  • add Matryoshka support

  • raise error

  • upd embed_dim in leaderboard

  • fix tests

  • fix typcheck

  • add tests

  • upd check

  • add docs (44e5947)

Unknown

Add a Thai language benchmark with 28 tasks spanning 6 task types:

  • BitextMining (6): BibleNLP, Flores, NTREX, Tatoeba, WebFAQ
  • Classification (9): Wisesight, Wongnai, SIB200, MASSIVE, MTOP, etc.
  • Clustering (1): SIB200ClusteringS2S
  • PairClassification (1): XNLI
  • Reranking (2): MIRACL, MultiLongDoc
  • Retrieval (9): MIRACL, BelebeleRetrieval, MKQARetrieval, etc.

Results for 13 models already merged: embeddings-benchmark/results#428

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

  • Update mteb/benchmarks/benchmarks/benchmarks.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Update benchmarks.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Curate Thai benchmark: remove cross-lingual and low-quality tasks

Address review feedback from @KennethEnevoldsen:

Removed (12 tasks):

  • All 6 bitext mining tasks — cross-lingual by design, not monolingual Thai
  • LanguageClassification — trivial LID, off-topic for monolingual benchmark
  • MassiveIntentClassification, MassiveScenarioClassification — machine-translated
  • MultilingualSentimentClassification — unclear Thai provenance
  • WongnaiReviewsClassification — redundant sentiment (keep Wisesight as best)

Kept (15 tasks):

  • Classification: MTOP (purpose-built for Thai), SIB200, Wisesight (native Thai)
  • Clustering: SIB200ClusteringS2S
  • PairClassification: XNLI (human-translated)
  • Reranking: MIRACL (human-judged), MultiLongDoc
  • Retrieval: MIRACL, Belebele, MKQA, MrTidy, MultiLongDoc, WebFAQ, XQuAD

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

  • Add contacts field for benchmark maintainer

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>


Co-authored-by: anusoft <anu@anusoft.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (03dcfc8)

  • model: Add NanoVDR-S-Multi with custom AbsEncoder for asymmetric VDR (#4242)

  • Add model meta for nanovdr/NanoVDR-S-Multi

  • add n_embedding_parameters

  • Implement NanoVDRWrapper as custom AbsEncoder with asymmetric routing

Query encoding uses the lightweight NanoVDR-S-Multi student (69M, text-only).
Document encoding uses the frozen Qwen3-VL-Embedding-2B teacher (2B, VLM).
Teacher is lazy-loaded only when document encoding is needed.

  • Fix ruff lint and formatting

  • Default to student encoder for non-retrieval tasks

  • Add error for unsupported image-query tasks

  • Remove trust_remote_code=True (model class built locally by MTEB)

  • Apply suggestion from @Samoed

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (2e1e513)