Skip to content

1.38.34

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 10 Jul 12:29
· 1579 commits to main since this release

1.38.34 (2025-07-10)

Fix

  • fix: pin datasets version (#2892)

fix datasets version (00c95cf)

Unknown

  • Update tasks & benchmarks tables (5303fec)

  • dataset: Evalita dataset integration (#2859)

  • Added DadoEvalCoarseClassification

  • Removed unnecessary columns from DadoEvalCoarseClassification

  • Added EmitClassification task

  • added SardiStanceClassification task

  • Added GeoLingItClassification task

  • Added DisCoTexPairClassification tasks

  • Added EmitClassification, DadoEvalCoarseClassification, GeoLingItClassification, SardiStanceClassification inside the inits

  • changed import in DisCoTexPairClassification

  • removed GeoLingItClassification dataset

  • fixed citation formatting, missing metadata parameters and lint formatting

    • Added XGlueWRPReranking task
  • Added missing init.py files
  • fixed metadata in XGlueWRPReranking

  • Added MKQARetrieval task

  • fixed type in XGlueWRPReranking

  • changed MKQARetrieval from cross-lingual to monolingual

  • formatted MKQARetrieval file

  • removed unused const


Co-authored-by: Mattia Sangermano <MattiaSangermano@users.noreply.huggingface.co> (ee17a6e)

  • model: add Hakim and TookaSBERTV2 models (#2826)

  • add tooka v2s

  • add mcinext models

  • update mcinext.py

  • Apply PR review suggestions

  • Update mteb/models/mcinext_models.py


Co-authored-by: mehran <mehan.sarmadi16@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (04dc6d4)

  • Update tasks & benchmarks tables (5be02c1)

  • Add and fix some Japanese datasets: ANLP datasets, JaCWIR, JQaRA (#2872)

  • Add JaCWIR and JQaRA for reranking

  • Fix ANLP Journal datasets

  • Add NLPJournalAbsArticleRetrieval and JaCWIRRetrieval

  • tackle test cases

  • Remove _evaluate_subset usage

  • Separate v1 and v2

  • Update info for NLP Journal datasets (70768b5)

  • Comment kalm model (#2877)

comment kalm model (a3ca95c)

  • model: add kalm_models ModelMeta (new PR) (#2853)

  • feat: add KaLM_Embedding_X_0605 in kalm_models

  • Update kalm_models.py for lint format


Co-authored-by: xinshuohu <xinshuohu@tencent.com> (b67bd04)

  • model: add listconranker modelmeta (#2874)

  • add listconranker modelmeta

  • fix bugs

  • use linter

  • lint


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (5846f56)

  • fix tests to be compatible with SentenceTransformers v5 (#2875)

  • fix sbert v5

  • add comment (f346a37)

  • rename seed-1.6-embedding to seed1.6-embedding (#2870) (f27648b)

  • model: Adding nvidia/llama-nemoretriever-colembed models (#2861)

  • nvidia_llama_nemoretriever_colembed

  • correct 3b reference

  • lint fix

  • add training data and license for nvidia/llama_nemoretriever_colembed

  • lint


Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (4ff1413)

  • Bump gradio to fix leaderboard sorting (#2866)

Bump gradio (a4388c2)

  • model: Adding Sailesh97/Hinvec (#2842)

  • Adding Hinvec Model's Meta data.

  • Adding hinvec_model.py

  • Update mteb/models/hinvec_models.py

Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>

  • formated code with Black and lint with Ruff

Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (e3286d5)