Repository navigation
2.20.8
2.20.8 (2026-09-05)
Fix
-
fix: repair broken Vultron revision pins (#5378)
-
fix: repair broken Vultron revision pins
-
fix: propagate ColQwen revision to all loaders
-
test: narrow ColQwen revision assertions
-
fix: update Vultron revisions
-
remove adapter_kwargs
-
lint
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (c9da0cb)
Unknown
-
dataset: add TAU Urban Acoustic Scenes 2022 clustering (a2a) (#5371)
-
dataset: add TAU Urban Acoustic Scenes 2022 clustering (a2a)
Closes #5019. mteb already scores this dataset with labels through
TAUAcousticScenes2022Mobile; this adds the unsupervised counterpart the issue
asks for, on the same stratified subsample at the same revision so the two are
directly comparable.
CLAP reaches v_measure 0.1680 against a 0.0049 random floor. That it does well
here while sitting at chance on speech retrieval is consistent with its training
on audio captions of environmental sound.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- cite the DCASE 2020 workshop paper, not just the Zenodo release
reference pointed at the dataset record, so the task had no paper behind it.
Mesaros, Heittola and Virtanen describe this recording collection and the
device-generalization setup in the DCASE 2020 workshop proceedings, which is
also what soundata cites for this dataset family. The Zenodo release stays in
the bibtex as the source of the actual files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (178c79f)
- dataset: add Vaani multilingual speech-text retrieval (45 languages) (#5352)
Adds VaaniA2TRetrieval and VaaniT2ARetrieval over 45 Indian languages, from
Vaani's transcribed release.
That release ships an official test split, so unlike the main Vaani corpus the
evaluation set is held out at source rather than sampled from training data.
45 of the 64 languages with a test split are kept; the rest fall under 25 usable
utterances. 4,001 clips.
Scripts were determined from the transcripts rather than assumed. Two are not
what the language name suggests: Chakma is romanised in this release despite
having its own script, and Tulu is written in Kannada.
Transcripts carry annotation markup - <noise>, <pause> and similar event tags
plus {...} braces marking code-switched English. Tags are removed as annotation
artefacts; braces are unwrapped so the code-switched word survives.
Byte-identical audio and repeated transcript text are dropped per language so
an identical query is not marked relevant to only one of the clips it matches.
main_score is hit_rate_at_5, matching every existing audio-text retrieval task
in mteb.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (dca9d97)