Skip to content

2.15.5

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 19 Jun 21:00
· 458 commits to main since this release

2.15.5 (2026-06-19)

Ci

  • ci: revert leaderboard refresh trusted publishing (#4831)

revert leaderboard refresh (14cf6d6)

  • ci: Setup hf trusted publishing (#4804)

  • setup hf trusted publishing

  • remove comment (a1f0a62)

Fix

  • fix: Get modalities of models from config (#4789)

infer modalities (2c17992)

Unknown

  • model: update VultronRetrieverPrime-Qwen3.5-8B repo path to the vultr org (#4836)

model: Update VultronRetrieverPrime-Qwen3.5-8B metadata to reflect new repository path (e091c0d)

  • Add VultronRetrieverPrime-Qwen3.5-8B ModelMeta (#4833)

Late-interaction (ColBERT MaxSim) visual document retriever: ColQwen3.5, dim 320,
8.4B params, Apache-2.0, 6 languages. Reuses the existing ColQwen3_5Wrapper.
Official ViDoRe scores V1 0.9208 / V2 0.6818 / V3 0.6472 (results PR to follow).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (99a830c)

  • Update Querit model implementation: Supplementing the citation information of Querit (#4824)

Update Querit model implementation: Supplementing the citation information of the Querit

Co-authored-by: zhongyunfei <zhongyunfei@baidu.com> (cfce4f2)

  • model: Add VIRTUE multimodal embedding models (Sony VIRTUE-2B/7B-SCaR) (#4822)

  • Add VIRTUE multimodal embedding models (Sony VIRTUE-2B/7B-SCaR)

  • address review feedback (41f6c7e)

  • Fix Querit model implementation: Supplementing the base model information of the Querit/Querit-4B (#4819)

  • Fix Querit model implementation: Supplementing the base model information of the Querit/Querit-4B

  • Fix Querit model implementation: Supplementing the base model information of the Querit/Querit-4B


Co-authored-by: zhongyunfei <zhongyunfei@baidu.com> (240da5b)

  • mveb: fix and unify domain tags across all 50 source datasets (#4738)

  • mveb: fix and unify domain tags across all 50 source datasets

The MVEB+ video task set had inconsistent and partially-wrong domains
tags. Issues fixed:

  • MSR-VTT had no domain tags at all (empty list). Now tagged ["Web"].
  • AVMeme-Exam was tagged with "Music" (it's internet memes, not music
    content). Now ["Entertainment", "Web"].
  • AudioCaps_AV was tagged "Encyclopaedic" (it's audio captioning). Now
    ["AudioScene", "Web"].
  • VGGSound was tagged just ["Web"] despite being audio-visual events.
    Now ["AudioScene", "Web"]. Same fix for VGGSound_AV_RETRIEVAL.
  • AV-SpeakerBench was tagged ("Web") on the base task and ("Spoken")
    on the PC variant --- same source data, inconsistent tags. Unified
    to ("Spoken").
  • WorldSense_1min was over-tagged with Entertainment+Music in some
    files and just ["Web"] in others. Unified to ["AudioScene", "Scene",
    "Web"].
  • Several datasets tagged "Spoken" without speech-driven content
    (DiDeMo, MSVD, ActivityNetCaptions, VATEX, panda-70m, TUNA-Bench).
    Removed the Spoken tag from those.
  • AVE-Dataset clustering tasks tagged with ["Music", "Scene", "Spoken"]
    (clearly wrong). Now aligned with the rest of AVE-Dataset:
    ["AudioScene", "Web"].
  • MELD was tagged just ["Entertainment"] across base and clustering
    variants; MELD is the Friends sitcom, so dialogue is central.
    Added "Spoken" -> ["Entertainment", "Spoken"].
  • UCF101 missing "Sport" tag. UCF101 has substantial Sport content.
    Now ["Scene", "Sport", "Web"].
  • Human-Animal-Cartoon missing "Entertainment" tag despite the cartoon
    domain. Now ["Entertainment", "Scene", "Web"].
  • PerceptionTest missing "Scene" tag despite being a scene-perception
    benchmark. Now ["Scene", "Web"].
  • Video-MME missing "Spoken" tag despite the narration-heavy content.
    Now ["Spoken", "Web"].
  • HMDB51 missing "Web" tag (sourced largely from web video). Now
    ["Scene", "Web"].
  • VideoCon, Vinoground (zachz/*) missing "Web" tag. Added.
  • RAVDESS tag list kept at ["Spoken"] (speech-emotion primary).
  • AVQA tag list extended with "AudioScene" (it's an audio-visual QA
    benchmark).

All 50 unique source datasets across 184 video tasks now have
consistent, non-empty domain tags. Verified by re-importing every
task: 184 tasks load cleanly.

Tags use only the existing TaskDomain Literal vocabulary in
task_metadata.py; no new domains added.

  • mveb: enrich domain taxonomy + fix mislabeled action/egocentric/meme datasets

Adds 5 video content domains to TaskDomain (Activity, Instructional,
Egocentric, Nature, Animation) and re-tags datasets that were mislabeled
or under-characterized, so the domain set actually reflects benchmark
content:

  • Action recognition (Kinetics-400/600/700, HMDB51, UCF101, SSv2,
    ActivityNet, VATEX, NExT-QA, Vinoground, VideoCon) -> Activity
    (was the catch-all "Scene", which means visual place/setting).
  • Breakfast, YouCook2 -> Instructional (cooking / how-to).
  • Diving48 -> Activity + Sport.
  • EgoSchema -> Egocentric (was bare "Web").
  • Human-Animal-Cartoon -> Activity + Animation + Nature.
  • AVMeme-Exam -> + Social (internet memes).
  • PerceptionTest -> drop misapplied "Scene".

Scene is now reserved for genuine visual-scene content (WorldSense).
All 184 video tasks load; every domain validates against TaskDomain.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>


Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (343df1a)

  • Update Querit model implementation: 4B version of Querit-Reranker newly open-sourced (#4808)

  • Update Querit model implementation: 4B version of Querit-Reranker newly open-sourced.

  • Update Querit model implementation: 4B version of Querit-Reranker newly open-sourced.

  • Update Querit model implementation: 4B version of Querit-Reranker newly open-sourced.

  • Apply suggestions from code review

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Apply suggestions from code review

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>


Co-authored-by: zhongyunfei <zhongyunfei@baidu.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (4976113)

  • Rename number_texts_intersect_with_train to samples_in_train (#4809) (e76f291)

  • polish MVEB leaderboard names + icons (#4803)

benchmarks: polish MVEB leaderboard names + icons

  • display_name: align scope names with MAEB/MIEB convention —
    "Video" -> "MVEB Video-Only" (cf. "MAEB Audio-Only"),
    "Text+Video" -> "MVEB Video-Text" (cf. MIEB "Image-Text").
    MVEB(beta) stays "MVEB" as the headline benchmark.
  • icon: switch the three displayed MVEB benchmarks from the monochrome,
    unpinned master/svg/libre-gui-activity icon to the colored,
    commit-pinned svg-color/libre-gui-video icon, matching the convention
    used by the other benchmarks.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (5039c18)