2.5.2
2.5.2 (2025-12-25)
Documentation
-
docs: update MIEB contributing guide for MTEB v2 AbsTask structure (#3787)
-
docs: update MIEB contributing guide for MTEB v2 AbsTask structure
-
Update docs/mieb/readme.md
-
Update docs/mieb/readme.md (
fb53f57)
Fix
-
fix: Add model_type in model_meta for all models (#3751)
-
Add model_type in model_meta for all models
-
added literal for model_type
-
update jina embedding model type
-
Added model_type to from_cross_encoder() method
-
update test
-
change location in model_meta to pass test
-
update late_interaction model and fix test
-
update late_interaction for colnomic models
-
update test
-
Update mteb/models/model_meta.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
-
fix naming
-
remove is_cross_encoder field and convert it into property
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (522eecc)
Unknown
-
Optimize validate filter scores only (#3792)
-
feat: add detailed timing logs to leaderboard initialization
Add comprehensive timing information to track performance of each step
in the leaderboard building process:
- Loading benchmark results (from cache or remote)
- Fetching and processing benchmarks
- Filtering models and generating tables
- Creating Gradio components and interface
- Prerun phase for cache population
Each step logs start and completion times with elapsed duration to help
identify performance bottlenecks during leaderboard initialization.
- perf: optimize benchmark processing with caching and vectorized operations
Implemented 3 high-impact optimizations to reduce benchmark processing time:
-
Cache get_model_metas() calls using @functools.lru_cache
- Eliminates 59 redundant calls (once per benchmark)
- Now called once and cached for all benchmarks
-
Replace pandas groupby().apply() with vectorized operations
- Replaced deprecated .apply(keep_best) pattern
- Uses sort_values() + groupby().first() instead
- Avoids nested function calls per group
-
Cache version string parsing with @functools.lru_cache
- Eliminates redundant parsing of same version strings
- Uses LRU cache with 10,000 entry limit
Performance improvements:
- Benchmark processing: 131.17s → 44.73s (2.93x faster, 66% reduction)
- join_revisions(): 84.96s → 1.73s (49x faster, 98% reduction)
- Leaderboard Step 3: 121.28s → 48.23s (2.51x faster, 60% reduction)
This significantly improves leaderboard startup time by reducing the
benchmark processing bottleneck.
-
perf: optimize validate_and_filter_scores filtering logic
-
Update mteb/results/task_result.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> (d3bc4cc)
-
Add leaderboard timing logs and join_revisions() speedups (#3790)
-
feat: add detailed timing logs to leaderboard initialization
Add comprehensive timing information to track performance of each step
in the leaderboard building process:
- Loading benchmark results (from cache or remote)
- Fetching and processing benchmarks
- Filtering models and generating tables
- Creating Gradio components and interface
- Prerun phase for cache population
Each step logs start and completion times with elapsed duration to help
identify performance bottlenecks during leaderboard initialization.
- perf: optimize benchmark processing with caching and vectorized operations
Implemented 3 high-impact optimizations to reduce benchmark processing time:
-
Cache get_model_metas() calls using @functools.lru_cache
- Eliminates 59 redundant calls (once per benchmark)
- Now called once and cached for all benchmarks
-
Replace pandas groupby().apply() with vectorized operations
- Replaced deprecated .apply(keep_best) pattern
- Uses sort_values() + groupby().first() instead
- Avoids nested function calls per group
-
Cache version string parsing with @functools.lru_cache
- Eliminates redundant parsing of same version strings
- Uses LRU cache with 10,000 entry limit
Performance improvements:
- Benchmark processing: 131.17s → 44.73s (2.93x faster, 66% reduction)
- join_revisions(): 84.96s → 1.73s (49x faster, 98% reduction)
- Leaderboard Step 3: 121.28s → 48.23s (2.51x faster, 60% reduction)
This significantly improves leaderboard startup time by reducing the
benchmark processing bottleneck.
- Update mteb/leaderboard/app.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
- fix: ensure deterministic revision grouping in join_revisions()
- Replace groupby(revision_clean) with groupby(revision)
- Remove non-deterministic iloc[0] access for revision selection
- Tasks with different original revisions (None vs external) now kept separate
- Each ModelResult has consistent revision across all its task_results
This resolves the issue where tasks with different original revisions that mapped
to the same cleaned value would be grouped together non-deterministically.
-
refactor: use default lru_cache maxsize for _get_cached_model_metas
-
refactor: remove optimization markers from comments
-
Apply suggestion from @isaac-chung
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> (57a4b0c)