Skip to content

2.12.26

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 21 Apr 23:47
· 731 commits to main since this release

2.12.26 (2026-04-21)

Fix

  • fix: HF benchmark result (#4344)

  • init benchmark eval results

  • add get score to benchmark

  • update scoring

  • add method for benchmark card creation

  • fix typing (990c1cf)

Unknown

  • [MVEB] Add fps implementation to Video Sampling (#4441)

  • fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification

  • fix: Move SIBFLEURS descriptive stats to AudioClassification

  • refactor: add FPS-based video frame sampling to collator

FramesCollator and VideoCollator now support two modes:

  • FPS-based (default, fps=2.0): frame count scales with video duration,
    with max_frames=128 as a safety cap for long videos
  • Fixed-sample (num_frames=N): always selects exactly N frames uniformly,
    preserving the previous behavior for models that need it

Existing callers (PE-AV, random baseline) switched to num_frames to
preserve their current fixed-sample behavior.

  • refactor: switch PE-AV and random baseline to FPS-based frame sampling

Both models now use the default FPS-based mode (fps=2.0, max_frames=256)
instead of fixed num_frames. This gives duration-proportional frame
coverage across videos of different lengths.

  • fix: address PR review - defaults to None, use end_stream_seconds
  • Set fps and max_frames defaults to None so models can skip collator
    resampling and let their own processors handle frame selection
  • Use video.metadata.end_stream_seconds for duration instead of
    computing num_frames / average_fps (handles VFR videos correctly)
  • When both fps and num_frames are None, return all frames as-is
    instead of raising an error
  • PE-AV and random baseline explicitly set fps=2.0 to avoid decoding
    all frames unnecessarily
  • refactor: expose collator params in PE-AV init

Allow fps, max_frames, num_frames, and max_samples to be configured
via the PE-AV wrapper constructor instead of being hardcoded.
Defaults to fps=2.0 matching the standard video understanding rate.

  • fix: address PR review - rename max_frames to max_fps_frames, raise on conflicting args

  • fix: rename max_fps_frames back to max_frames, clarify docstrings

  • fix: pass fps=None, num_frames=16 to 16-frame PE-AV variants

The *-16-frame checkpoints were trained with fixed 16-frame uniform sampling
(processor config has do_sample_frames=true, num_frames=16). Without
explicit loader_kwargs, the collator used the default fps=2.0, producing
~40 frames on typical clips that the processor then re-sampled down to 16 —
a distribution shift from training. Setting num_frames=16 makes the
collator do the sampling directly, and the processor's built-in sample
becomes an identity no-op.

  • fix: clarify fps docstrings - downsamples only, no upsampling (9363ea7)

  • Don't display license links in the documentation (#4465)

Fixes #4461 (46582d9)

  • [MVEB] Adding UCF101 Task (Clustering) (#4454) (b8b3722)

  • leaderboard: add MTEB(spa, v1) to Language-specific section (#4217)

Add MTEB(spa, v1) to leaderboard language-specific menu

Co-authored-by: Clemente <clemente@Clementes-MacBook-Pro.local> (e5521a6)

  • Add VALOR-32K retrieval tasks (#4453)

  • Add VALOR-32K retrieval tasks (v2t, t2v, va2t, t2va)

Adds four bidirectional multimodal retrieval tasks for the VALOR-32K
dataset (mteb/VALOR-32K), a vision-audio-language benchmark with 3,491
test samples.

Made-with: Cursor

  • fix: correct BibTeX field order for VALOR-32K citation

Made-with: Cursor (792f61f)