Skip to content

2.21.6

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 21 Sep 22:39
· 47 commits to main since this release

2.21.6 (2026-09-21)

Fix

  • fix: remaining changes before making import mteb torch-free and mteb-core compatible (#5502)

  • fix: remaining changes before making import mteb torch-free and mteb-core compatible

  • changes from review

  • fix: clean up stale comment and dead code from torch-free import refactor

Update a comment in test_ensure_no_torch_at_import.py that claimed
import mteb still pulls in torch, contradicted by
test_import_mteb_does_not_import_torch further down in the same file.
Also remove the now-unused logger in _set_seed.py left over after the
tensorflow debug-log call was dropped.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>


Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> (e07c2e3)

Unknown

  • dataset: Add chinaopen multilingual video (#5393)

  • task: add ChinaOpen multilingual video retrieval (t2v, v2t)

Adds the first multilingual video retrieval tasks in mteb, built from
ChinaOpen-1k (ACM MM 2023): 1,092 Bilibili videos whose Chinese captions
are written by human annotators and whose English captions are
translations of them. Both language subsets describe the same videos, so
zho-Hans vs eng-Latn is a controlled comparison rather than a difference
in visual content.

  • ChinaOpenT2VRetrieval: caption -> video
  • ChinaOpenV2TRetrieval: video -> caption

Queries are deduplicated by caption text and every video carrying a
caption is marked relevant, so the few captions shared by more than one
video are multi-positive instead of being scored as misses. Uploader
video titles are not used, only the human-written manual captions.

The datasets are published in the standard retrieval layout with one
config triple per language, so no custom load_data is needed. The
construction script is included.

  • test: allow shared videos across ChinaOpen language subsets

Both language subsets of the ChinaOpen tasks describe the same 1,092
videos, which is the point of the controlled comparison, so the corpus
is intentionally shared and the duplicate check counts it twice. This
is the same situation as XM3600 and XFlickr30k-Co, which are listed
under duplicate_image for the same reason.

  • refactor: combine ChinaOpen T2V/V2T into a single shared-video dataset

Rebuild the ChinaOpen dataset as one HF repo (mteb/ChinaOpen1k) with a
shared videos config plus per-language &lt;lang&gt;-texts/&lt;lang&gt;-links
configs, instead of two full repos that each duplicated the video bytes
per direction. ChinaOpenT2VRetrieval and ChinaOpenV2TRetrieval now share
a _load_chinaopen(task, direction) loader that swaps which side of each
caption<->video link is the query, keeping the original dedup and
multi-positive qrels logic intact.

Verified end-to-end: rebuilt from the raw ChinaOpen-1k source, pushed to
mteb/ChinaOpen1k, and reran both tasks with the random-encoder baseline,
reproducing the original PR's reference ndcg@10 scores.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

  • chore: regenerate ChinaOpen descriptive stats against combined dataset

Recompute descriptive_stats JSON for ChinaOpenT2VRetrieval and
ChinaOpenV2TRetrieval against the new mteb/ChinaOpen1k dataset via
calculate_descriptive_statistics(overwrite_results=True). Values match
the original PR's stats exactly (num_samples=4360, num_queries=2176/2184,
num_documents=2184/2176), confirming the restructuring didn't change the
underlying data. Also includes a ruff-format pass on create_data.py.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>


Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> (52531fc)