Repository navigation
2.20.5
2.20.5 (2026-08-31)
Fix
-
fix: add recording-level FleursT2ARetrieval.v2 / FleursA2TRetrieval.v2 (#5307)
-
fix: add recording-level FleursT2ARetrieval.v2 / FleursA2TRetrieval.v2
FLEURS id identifies a sentence, not a recording: each sentence is read by
up to six different native speakers who all share one id. Both Fleurs
retrieval tasks key queries, corpus and qrels on that id, so all recordings of
a sentence collapse into a single id. They are not dropped -- every row is
still encoded and scored -- but for T2A they merge under one result key, so
aggregation over speakers happens implicitly and depends on row order rather
than being declared as multi-positive qrels; for A2T they are never evaluated
as independent queries. Across the 102 subsets 77,809 recordings carry only
33,018 distinct ids.
The FLEURS paper defines both directions explicitly (Conneau et al., SLT 2022,
Sec. 4): speech-to-text retrieves "the correct text segment", and text-to-speech
scores "retrieving any of the speakers who speaks the correct textual query".
That is exactly hit_rate_at_k (success@k), which both tasks already use as
main_score -- the metric was right, only the qrels were not.
v2 matches those semantics:
- T2A.v2 -- one query per unique sentence; every physical recording gets a
unique corpus id; all recordings of a sentence are positives
(average_relevant_docs_per_query 2.357, max 6). - A2T.v2 -- every recording is an independent query; the corpus holds one
document per unique sentence rather than duplicated identical transcripts.
Recording ids are {sentence_id}-{rank}, ranked by the globally unique FLEURS
audio filename, so an id is invariant to row order. In ln_cd two distinct
sentence ids carry byte-identical text; those documents are indistinguishable,
so both are marked positive rather than merging the ids.
v1 is left intact so published results stay valid; it only gains a
superseded_by pointer, following the BSARDRetrieval -> BSARDRetrieval.v2
convention for a one-to-many qrel fix. get_tasks(tasks=[...]) short-circuits
before exclude_superseded, so MAEB(beta) still resolves v1 by name.
Descriptive stats for the new tasks reuse v1's audio statistics, justified by
verifying the audio column is byte-identical between the v1 and v2
constructions; the text and qrel statistics are recomputed. Truncated audio in
the upstream data is a separate data-quality issue and is deliberately not
filtered here.
Fixes #5270
-
Delete tests/test_tasks/test_fleurs_v2.py
-
lint
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (ef85903)
Unknown
- dataset: add WIT image-to-text multilingual retrieval (#5326)
dataset: add WIT image-to-text retrieval (0d7a65c)
-
dataset: add XM3600 image-to-text multilingual retrieval (#5325)
-
dataset: add XM3600 image-to-text retrieval
-
dataset: address XM3600 I2T review feedback (
05a344c) -
model: Pass kwargs to bm25s (#5268)
pass kwargs to bm25s (9e96b20)
-
dataset: add XFlickr30k-Co image-to-text retrieval (#5324)
-
dataset: add XFlickr30k-Co image-to-text retrieval
-
fix: address XFlickr30k-Co review feedback (
2fa7131) -
Enable
A(flake8-builtins) ruff rule (#5328)
enable A ruff rule (d40e259)
- dataset: add DROID it2v robot manipulation retrieval (MOEB) (#5256)
Adds DROIDIT2VRetrieval, composed image+text -> video retrieval over the
DROID in-the-wild Franka manipulation dataset. 1,500 successful episodes
with unique language instructions (5-60 s at 15 fps), evenly sampled
across the collection. The query pairs the initial scene image from the
exterior_1 camera with the instruction; the corpus holds exterior_2
videos of the same episodes, so queries and documents never share a
viewpoint and exact frame matching cannot solve the task. Relevance is
instance-level 1:1.
Reference runs (ndcg@10): mteb/baseline-random-encoder 0.0028;
jinaai/jina-embeddings-v5-omni-nano 0.1691.
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (86079bf)
-
dataset: add BridgeData V2 v2v robot manipulation retrieval (MOEB) (#5259)
-
dataset: add BridgeData V2 v2v robot manipulation retrieval (MOEB)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
- lint: format bridge construction script
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
- fix: correct BridgeData V2 license to cc-by-4.0
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (f1a3c51)