Skip to content

2.21.1

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 20 Sep 21:37
· 54 commits to main since this release

2.21.1 (2026-09-20)

Fix

  • fix: make model implementations importable without torch (#5463)

  • fix: make model implementations importable without torch

  • add gme, mixedbread and semantic router models changes

  • add mps in get_device

  • updated docs

  • lintter

  • changes from review

  • updated docs (88058a3)

  • fix: warm the benchmark-schema cache key the route actually uses (#5494)

  • fix: warm the benchmark-schema cache key the route actually uses

_prewarm_list_schemas submitted _benchmark_schemas_bytes with no arguments,
but /v1/benchmarks always calls it positionally. functools.lru_cache keys on
the argument tuple, so the warmed () entry is never read and the first request
rebuilds the list itself.

  • Delete tests/test_api_warmup.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (21f8148)

  • fix: pass num_proc to every dataloader created during evaluation (#5489)

  • fix: pass num_proc to every dataloader created during evaluation

  • simplify test


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (322d09f)

  • fix: apply scripts filter in TaskResult.get_score (#5490)

  • fix: apply scripts filter in TaskResult.get_score

  • Apply suggestion from @Samoed

  • lint


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (9de9d1c)

  • fix: accept ISO 639-3 codes in get_model_metas(languages=...) (#5484)

  • fix: accept ISO 639-3 codes in get_model_metas(languages=...)

  • simplify

  • simplify


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (8ecdd97)

  • fix: accept language-script and programming-language codes in filter_tasks (#5483)

  • fix: accept language-script and programming-language codes in filter_tasks

  • simplify check

  • simplify check


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (67bd1b8)

  • fix: allow validate_and_filter in load_results without tasks (#5480) (af2708f)

  • fix: do not mutate eval_splits in calculate_descriptive_statistics (#5481) (c42a4e6)

  • fix: pass the selected benchmarks in create-model-results (#5477) (f9b5cf2)

  • fix: fix citations (#5473)

fix citations (0d9d487)

  • fix: apply the n_parameters lower bound when there is no upper bound (#5476)

The lower-bound check was nested inside 'if upper is not None', so a
(lower, None) range filtered nothing. Hoist the guard so either bound
activates the filter; upper-only behaviour is unchanged. (629aa89)

  • fix: Correct citations for SpokenSQuAD and CMU (#5472)

Addresses issues mentioned in #5471 (closing only once we have added tests) (952bb7f)

  • fix: reopen existing NumpyCache without a -1 memmap shape (#5470)

np.memmap does not infer a -1 dimension, so NumpyCache.load() failed on every
existing cache directory (OverflowError on numpy<2.2, ValueError on numpy>=2.2)
and CachedEmbeddingWrapper could never reuse embeddings from an earlier run.
Derive the row count from the file size instead. (00407be)

Refactor

  • refactor: move FreshStackRetrieval to v2 dataset format (#5495) (840b824)

Unknown

  • model: add albertobarnabo/bge-m3-italian (#5488)

  • model: add albertobarnabo/bge-m3-italian

  • model: drop HF-only citation for bge-m3-italian (ee57f4e)

  • dataset: add MM-BRIGHT retrieval tasks (#5168)

  • dataset: add MM-BRIGHT reranking tasks

  • dataset: defer Pillow import for MM-BRIGHT

  • reupload


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (5d3212e)

  • Delete results directory (#5485) (f32c71b)

  • Add NQ-Tables retrieval (#5461)

  • Add unregistered NQ-Tables retrieval engineering prototype

  • Remove NQ-Tables implementation notes from prototype PR

  • Remove dangling qrel explanatory comment

  • Add test-only NQ-Tables descriptive statistics

  • Remove NQ-Tables validation artifacts from task PR

  • Register NQ-Tables as a standard retrieval task

  • Promote NQ-Tables from prototype status

  • Address NQ-Tables review feedback

  • format citation


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (a0ee007)

  • dataset: add JuriFindITRetrieval (#5464)

  • dataset: add JuriFindITRetrieval

  • dataset: build a single test split for JuriFindITRetrieval (29b21ba)

  • fix typing (#5474) (f4ed2de)

  • model: restore bidirectional-attention fallback for PyPI colpali_engine (#5469)

  • fix(evie): restore bidirectional-attention fallback for PyPI colpali_engine

ColQwen3_5.enable_bidirectional_attention ships with the EVIE release of
colpali_engine but is absent from the published PyPI wheels that the evie
extra resolves to, so EvieWrapper.init raised AttributeError on a real
load (mteb#5451). Mock tasks never call from_pretrained, so CI missed it.

Restore the guard with an equivalent local implementation. Both the config
and the module flag have to be flipped: config.is_causal is what
create_causal_mask reads (and what sdpa honours on padded batches), while
Qwen3_5Attention.is_causal is what flash_attention_2 reads. Flipping only
the module flag leaves sdpa silently causal.

Verified bit-exact against the released implementation on padded batches
under both sdpa and flash_attention_2.

  • review: link the reference implementation, drop redundant comment

  • lint


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (c9f5262)

  • Add model: Singaraj/sante-embed (#5466)

  • Add model: Singaraj/sante-embed

  • Delete mteb_mock_run_results.md


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (8765064)

  • model: Add NGA-KR/ja-embed-sts ModelMeta (#5467)

Add NGA-KR/ja-embed-sts ModelMeta (46003a1)

  • dataset: add NayanaIR monolingual visual document retrieval (t2i, 8 Indic languages) (#5458)

  • dataset: add NayanaIR monolingual visual document retrieval (t2i, 8 Indic languages)

  • Update mteb/tasks/retrieval/multilingual/nayanair_monobench_retrieval.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>


Co-authored-by: gitgod-debug <186448411+gitgod-debug@users.noreply.github.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (fc92859)

  • Create aurola_omni_models (#5453)

  • Create aurola_omni_models.py

Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>

  • update license

Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>

  • Simplify code

Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>


Co-authored-by: nehal309 <20415842+nehal309@users.noreply.github.com> (3ab2623)