Skip to content

2.5.2

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 25 Dec 23:21
· 1165 commits to main since this release

2.5.2 (2025-12-25)

Documentation

  • docs: update MIEB contributing guide for MTEB v2 AbsTask structure (#3787)

  • docs: update MIEB contributing guide for MTEB v2 AbsTask structure

  • Update docs/mieb/readme.md

  • Update docs/mieb/readme.md (fb53f57)

Fix

  • fix: Add model_type in model_meta for all models (#3751)

  • Add model_type in model_meta for all models

  • added literal for model_type

  • update jina embedding model type

  • Added model_type to from_cross_encoder() method

  • update test

  • change location in model_meta to pass test

  • update late_interaction model and fix test

  • update late_interaction for colnomic models

  • update test

  • Update mteb/models/model_meta.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • fix naming

  • remove is_cross_encoder field and convert it into property


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (522eecc)

Unknown

  • Optimize validate filter scores only (#3792)

  • feat: add detailed timing logs to leaderboard initialization

Add comprehensive timing information to track performance of each step
in the leaderboard building process:

  • Loading benchmark results (from cache or remote)
  • Fetching and processing benchmarks
  • Filtering models and generating tables
  • Creating Gradio components and interface
  • Prerun phase for cache population

Each step logs start and completion times with elapsed duration to help
identify performance bottlenecks during leaderboard initialization.

  • perf: optimize benchmark processing with caching and vectorized operations

Implemented 3 high-impact optimizations to reduce benchmark processing time:

  1. Cache get_model_metas() calls using @functools.lru_cache

    • Eliminates 59 redundant calls (once per benchmark)
    • Now called once and cached for all benchmarks
  2. Replace pandas groupby().apply() with vectorized operations

    • Replaced deprecated .apply(keep_best) pattern
    • Uses sort_values() + groupby().first() instead
    • Avoids nested function calls per group
  3. Cache version string parsing with @functools.lru_cache

    • Eliminates redundant parsing of same version strings
    • Uses LRU cache with 10,000 entry limit

Performance improvements:

  • Benchmark processing: 131.17s → 44.73s (2.93x faster, 66% reduction)
  • join_revisions(): 84.96s → 1.73s (49x faster, 98% reduction)
  • Leaderboard Step 3: 121.28s → 48.23s (2.51x faster, 60% reduction)

This significantly improves leaderboard startup time by reducing the
benchmark processing bottleneck.

  • perf: optimize validate_and_filter_scores filtering logic

  • Update mteb/results/task_result.py

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>


Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> (d3bc4cc)

  • Add leaderboard timing logs and join_revisions() speedups (#3790)

  • feat: add detailed timing logs to leaderboard initialization

Add comprehensive timing information to track performance of each step
in the leaderboard building process:

  • Loading benchmark results (from cache or remote)
  • Fetching and processing benchmarks
  • Filtering models and generating tables
  • Creating Gradio components and interface
  • Prerun phase for cache population

Each step logs start and completion times with elapsed duration to help
identify performance bottlenecks during leaderboard initialization.

  • perf: optimize benchmark processing with caching and vectorized operations

Implemented 3 high-impact optimizations to reduce benchmark processing time:

  1. Cache get_model_metas() calls using @functools.lru_cache

    • Eliminates 59 redundant calls (once per benchmark)
    • Now called once and cached for all benchmarks
  2. Replace pandas groupby().apply() with vectorized operations

    • Replaced deprecated .apply(keep_best) pattern
    • Uses sort_values() + groupby().first() instead
    • Avoids nested function calls per group
  3. Cache version string parsing with @functools.lru_cache

    • Eliminates redundant parsing of same version strings
    • Uses LRU cache with 10,000 entry limit

Performance improvements:

  • Benchmark processing: 131.17s → 44.73s (2.93x faster, 66% reduction)
  • join_revisions(): 84.96s → 1.73s (49x faster, 98% reduction)
  • Leaderboard Step 3: 121.28s → 48.23s (2.51x faster, 60% reduction)

This significantly improves leaderboard startup time by reducing the
benchmark processing bottleneck.

  • Update mteb/leaderboard/app.py

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

  • fix: ensure deterministic revision grouping in join_revisions()
  • Replace groupby(revision_clean) with groupby(revision)
  • Remove non-deterministic iloc[0] access for revision selection
  • Tasks with different original revisions (None vs external) now kept separate
  • Each ModelResult has consistent revision across all its task_results

This resolves the issue where tasks with different original revisions that mapped
to the same cleaned value would be grouped together non-deterministically.

  • refactor: use default lru_cache maxsize for _get_cached_model_metas

  • refactor: remove optimization markers from comments

  • Apply suggestion from @isaac-chung


Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> (57a4b0c)

  • model: add octen_models (#3789)

  • model: add octen_models

  • add issue link for document prompt (a2631dd)

  • better clustering fix (#3793) (44555a7)