2.15.5
2.15.5 (2026-06-19)
Ci
- ci: revert leaderboard refresh trusted publishing (#4831)
revert leaderboard refresh (14cf6d6)
Fix
- fix: Get modalities of models from config (#4789)
infer modalities (2c17992)
Unknown
- model: update VultronRetrieverPrime-Qwen3.5-8B repo path to the vultr org (#4836)
model: Update VultronRetrieverPrime-Qwen3.5-8B metadata to reflect new repository path (e091c0d)
- Add VultronRetrieverPrime-Qwen3.5-8B ModelMeta (#4833)
Late-interaction (ColBERT MaxSim) visual document retriever: ColQwen3.5, dim 320,
8.4B params, Apache-2.0, 6 languages. Reuses the existing ColQwen3_5Wrapper.
Official ViDoRe scores V1 0.9208 / V2 0.6818 / V3 0.6472 (results PR to follow).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (99a830c)
- Update Querit model implementation: Supplementing the citation information of Querit (#4824)
Update Querit model implementation: Supplementing the citation information of the Querit
Co-authored-by: zhongyunfei <zhongyunfei@baidu.com> (cfce4f2)
-
model: Add VIRTUE multimodal embedding models (Sony VIRTUE-2B/7B-SCaR) (#4822)
-
Add VIRTUE multimodal embedding models (Sony VIRTUE-2B/7B-SCaR)
-
address review feedback (
41f6c7e) -
Fix Querit model implementation: Supplementing the base model information of the Querit/Querit-4B (#4819)
-
Fix Querit model implementation: Supplementing the base model information of the Querit/Querit-4B
-
Fix Querit model implementation: Supplementing the base model information of the Querit/Querit-4B
Co-authored-by: zhongyunfei <zhongyunfei@baidu.com> (240da5b)
-
mveb: fix and unify domain tags across all 50 source datasets (#4738)
-
mveb: fix and unify domain tags across all 50 source datasets
The MVEB+ video task set had inconsistent and partially-wrong domains
tags. Issues fixed:
- MSR-VTT had no domain tags at all (empty list). Now tagged ["Web"].
- AVMeme-Exam was tagged with "Music" (it's internet memes, not music
content). Now ["Entertainment", "Web"]. - AudioCaps_AV was tagged "Encyclopaedic" (it's audio captioning). Now
["AudioScene", "Web"]. - VGGSound was tagged just ["Web"] despite being audio-visual events.
Now ["AudioScene", "Web"]. Same fix for VGGSound_AV_RETRIEVAL. - AV-SpeakerBench was tagged ("Web") on the base task and ("Spoken")
on the PC variant --- same source data, inconsistent tags. Unified
to ("Spoken"). - WorldSense_1min was over-tagged with Entertainment+Music in some
files and just ["Web"] in others. Unified to ["AudioScene", "Scene",
"Web"]. - Several datasets tagged "Spoken" without speech-driven content
(DiDeMo, MSVD, ActivityNetCaptions, VATEX, panda-70m, TUNA-Bench).
Removed the Spoken tag from those. - AVE-Dataset clustering tasks tagged with ["Music", "Scene", "Spoken"]
(clearly wrong). Now aligned with the rest of AVE-Dataset:
["AudioScene", "Web"]. - MELD was tagged just ["Entertainment"] across base and clustering
variants; MELD is the Friends sitcom, so dialogue is central.
Added "Spoken" -> ["Entertainment", "Spoken"]. - UCF101 missing "Sport" tag. UCF101 has substantial Sport content.
Now ["Scene", "Sport", "Web"]. - Human-Animal-Cartoon missing "Entertainment" tag despite the cartoon
domain. Now ["Entertainment", "Scene", "Web"]. - PerceptionTest missing "Scene" tag despite being a scene-perception
benchmark. Now ["Scene", "Web"]. - Video-MME missing "Spoken" tag despite the narration-heavy content.
Now ["Spoken", "Web"]. - HMDB51 missing "Web" tag (sourced largely from web video). Now
["Scene", "Web"]. - VideoCon, Vinoground (zachz/*) missing "Web" tag. Added.
- RAVDESS tag list kept at ["Spoken"] (speech-emotion primary).
- AVQA tag list extended with "AudioScene" (it's an audio-visual QA
benchmark).
All 50 unique source datasets across 184 video tasks now have
consistent, non-empty domain tags. Verified by re-importing every
task: 184 tasks load cleanly.
Tags use only the existing TaskDomain Literal vocabulary in
task_metadata.py; no new domains added.
- mveb: enrich domain taxonomy + fix mislabeled action/egocentric/meme datasets
Adds 5 video content domains to TaskDomain (Activity, Instructional,
Egocentric, Nature, Animation) and re-tags datasets that were mislabeled
or under-characterized, so the domain set actually reflects benchmark
content:
- Action recognition (Kinetics-400/600/700, HMDB51, UCF101, SSv2,
ActivityNet, VATEX, NExT-QA, Vinoground, VideoCon) -> Activity
(was the catch-all "Scene", which means visual place/setting). - Breakfast, YouCook2 -> Instructional (cooking / how-to).
- Diving48 -> Activity + Sport.
- EgoSchema -> Egocentric (was bare "Web").
- Human-Animal-Cartoon -> Activity + Animation + Nature.
- AVMeme-Exam -> + Social (internet memes).
- PerceptionTest -> drop misapplied "Scene".
Scene is now reserved for genuine visual-scene content (WorldSense).
All 184 video tasks load; every domain validates against TaskDomain.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (343df1a)
-
Update Querit model implementation: 4B version of Querit-Reranker newly open-sourced (#4808)
-
Update Querit model implementation: 4B version of Querit-Reranker newly open-sourced.
-
Update Querit model implementation: 4B version of Querit-Reranker newly open-sourced.
-
Update Querit model implementation: 4B version of Querit-Reranker newly open-sourced.
-
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
- Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: zhongyunfei <zhongyunfei@baidu.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (4976113)
-
Rename
number_texts_intersect_with_traintosamples_in_train(#4809) (e76f291) -
polish MVEB leaderboard names + icons (#4803)
benchmarks: polish MVEB leaderboard names + icons
- display_name: align scope names with MAEB/MIEB convention —
"Video" -> "MVEB Video-Only" (cf. "MAEB Audio-Only"),
"Text+Video" -> "MVEB Video-Text" (cf. MIEB "Image-Text").
MVEB(beta) stays "MVEB" as the headline benchmark. - icon: switch the three displayed MVEB benchmarks from the monochrome,
unpinned master/svg/libre-gui-activity icon to the colored,
commit-pinned svg-color/libre-gui-video icon, matching the convention
used by the other benchmarks.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (5039c18)