Skip to content

2.5.0

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 23 Dec 13:10
· 1174 commits to main since this release

2.5.0 (2025-12-23)

Feature

  • feat: Added get_benchmark_result() to BenchmarkResults to obtain a benchmark table (#3771)

  • Update BenchmarkResults to output results of benchmark

  • added score column and correct TYPE_CHECKING

  • address comments

  • address comments

  • fix import

  • fix tests

  • fix tests

  • change BenchmarkResults to Pydantic dataclass

  • change benchmark to pydantic dataclass

  • fix tests

  • fix model

  • fix

  • lint

  • remove future

  • fix after review

  • add test

  • reapply comments from review

  • remove mock benchmark

  • add documentation

  • added actual results

  • Update docs/usage/loading_results.md

Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>

  • add actual results

Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (92fa705)

Unknown

  • Updates on ColEmbed 3B and 1B model wrapper to support MTEB 2 image dataloader (#3783)

Renaming model forward_passages to forward_images (0a0e398)

  • dataset: Add Turkish Constitutional Court violation classification task (#3777)

  • Add Turkish Constitutional Court violation classification task

  • Format Turkish task files with ruff

  • Update mteb/tasks/classification/tur/turkish_constitutional_court.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Fix dataset revision to pinned commit hash

  • Remove results directory from task PR

  • task dataset path updated

  • load data deleted

  • ruff formatted

  • bibtex fix

  • descriptive stats added

  • descriptive stats file name fix

  • Fix dataset duplicates/overlap and bibtex

  • Fix dataset duplicates, remove validation, and clean BibTeX

  • Format TurkishConstitutionalCourtViolation task with ruff

  • upd statistics


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (d3ce3d5)

  • model: add c2llm & fix f2llm (#3782)

  • add c2llm & fix f2llm

  • Update c2llm via sentence_transformer_wrapper

  • add c2llm languages

  • format


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (9b00964)

  • dataset: add SQuADKorV1Retrieval task for Korean (#3779)

  • dataset: add SQuADKorV1Retrieval task for Korean

Add Korean SQuAD v1.0 retrieval task based on the yjoonjang/squad_kor_v1 dataset.

This dataset provides:

  • 5,777 queries
  • 960 corpus documents
  • Korean Wikipedia-based question answering pairs
  • dataset: add descriptive statistics for SQuADKorV1Retrieval

Statistics:

  • 6,734 total samples
  • 960 unique documents (avg 545 chars)
  • 5,764 unique queries (avg 34 chars) (fd339dd)
  • Fix: update reference website of Seed1.6-embedding-1215 (#3780)

update reference website of Seed1.6-embedding-1215 (f5182fb)

  • fix pre commit (#3775) (1a7b4cf)

  • fix_mod_embedding_oom (#3774) (02bdcd8)

  • model: added Bytedance/Seed1.6-embedding-1215 (#3760)

  • add model: Bytedance/Seed1.6-embedding-1215

  • make lint

  • fix typo

  • remove get_text_embedding() and get_image_embedding(), only reserve get_fused

  • move "max_tokens" and "available_embed_dims" into model implementations

  • fix typo

  • Apply suggestions from code review

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • fix: fix bug when "images" or "texts" are none

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (0d73c3d)

  • Added revision parameter for results repo (#3764)

  • Added revision parameter for results repo

  • fix clone_cmd

  • Added depth parameter in clone command (e8c02b1)