Skip to content

2.3.1

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 03 Dec 17:02
· 1253 commits to main since this release

2.3.1 (2025-12-03)

Fix

  • fix: colpali_training_set & updated JinaVDR and ViDoRe tasks annotation (#3636)

  • add adapted annotation

  • fix training set annotation

  • update nemotriever datasets (1ce74c2)

  • fix: add flag to run public only tasks (#3563)

  • run public only tasks

  • pass to evaluate

  • add tests

  • update cli

  • add task error

  • fix metadata name

  • fix tests

  • Update mteb/evaluate.py

Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>

  • remove from cli

  • add tests

  • fix exception

  • fix renaming

  • raise error if public_only False

  • rollback co2

  • Apply suggestions from code review

Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>

  • Apply suggestions from code review

Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>

  • remove test

  • fix test


Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (f905b68)

Unknown

  • model: Add IEITYuan/Yuan-embedding-2.0-en model (#3630)

  • add

  • Update mteb/models/model_implementations/yuan_models_en.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Update mteb/models/model_implementations/yuan_models_en.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Update mteb/models/model_implementations/yuan_models_en.py

Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>

  • Update mteb/models/model_implementations/yuan_models_en.py

Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>

  • Apply suggestions from code review

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (d31f1cd)

  • model: Add tomoro-colqwen3-embed embedding models (#3627)

  • feat(colqwen3): add wrapper and model metadata

  • feat(colqwen3): update ColQwen3Wrapper to use bfloat16 and enhance similarity scoring

  • Changed default dtype from float16 to bfloat16 for improved performance.
  • Added max_num_visual_tokens parameter to AutoProcessor initialization.
  • Refined embedding extraction logic to avoid boolean casting issues.
  • Introduced support for score_multi_vector in similarity computation.
  • Added new model metadata for colqwen3_4b with relevant attributes.
  • fix(colqwen): require transformers>=4.57 and refresh metadata, set revision

  • refactor(colqwen): reorder wrappers and metadata definitions for clarity

  • chore(colqwen): set release date for tomoro colqwen3 8b

  • chore(colqwen): remove unused methods and fix lint errors

  • feat(colqwen3): add fused image-text encoding path

  • refactor(colqwen): unify encode method with get_fused_embeddings

  • chore(colqwen): update encoding progress message

  • chore(colqwen): update model revisions for colqwen models

  • docs(colqwen): update train data annotation

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>


Co-authored-by: tankm <kyemin.tan@tomoro.ai>
Co-authored-by: Huang Xin <hxssg1124@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (71ac96c)

  • model: euler legal embedding (#3640)

  • add model implementation

  • modify the correct model path

  • Update mteb/models/model_implementations/euler_models.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Update mteb/models/model_implementations/euler_models.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Update mteb/models/model_implementations/euler_models.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • Update mteb/models/model_implementations/euler_models.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • lint

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (a1fbdb9)

  • model: jina-reranker-v3 (#3645)

  • add: jina-reranker-v3

  • Update mteb/models/model_implementations/jina_models.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • fix: remove useless check

  • fix: use only one class which inherits from CrossEncoderWrapper

  • Update mteb/models/model_implementations/jina_models.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • fix: remove unused import

  • Update mteb/models/model_implementations/jina_models.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • fix: ruff format

  • add: cross_encoder label

  • add: datasets

  • Update mteb/models/model_implementations/jina_models.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (1ecd892)

  • dataset: MultiLongDocReranking (#3642)

  • Add MultiLongDocReranking dataset

  • Add descriptive stats for MultiLongDocReranking

  • Fix information

  • reformat

  • fix hash

  • Update multi_long_doc_reranking.py

  • lint


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4f7774d)