Skip to content

2.12.16

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 10 Apr 14:46
· 796 commits to main since this release

2.12.16 (2026-04-10)

Fix

  • fix: add language extraction from HF model cards to get_model_meta (#4278)

  • feat: add language extraction from HF model cards to get_model_meta

Extract language codes from HuggingFace model card metadata and convert
them to MTEB's internal ISOLanguageScript format (e.g., "eng-Latn").

  • Add ISO 639-1 (2-letter) to ISO 639-3 (3-letter) mapping (183 codes)
  • Add default script mapping from ISO 639-3 to ISO 15924 (303 codes)
  • Add hf_lang_to_iso_lang_script() and hf_langs_to_iso_lang_scripts()
    conversion functions
  • Wire language extraction into ModelMeta._from_hub()
  • Add ISO language code false positives to typos config

Closes #3694

  • fix: remove unnecessary typos extend-words for language codes

The language codes (ful, som, yor, etc.) only appear in JSON files
which are already excluded by [tool.typos.files] extend-exclude.

  • fix: lazy-load JSON mappings and make HF lang functions private

Address KennethEnevoldsen's review comments:

  • Lazy-load iso_639_1_to_3.json and iso_639_3_to_default_script.json
    to avoid doing too much at import time
  • Rename hf_lang_to_iso_lang_script -> _hf_lang_to_iso_lang_script
    and hf_langs_to_iso_lang_scripts -> _hf_langs_to_iso_lang_scripts
    (not a public interface we want to maintain)
  • fix: use functools.cache for lazy loading instead of global statement

Fixes PLW0603 lint error (global statement discouraged). The private
function names are kept per KennethEnevoldsen's review.

  • fix: resolve CI lint errors (PLR0911, PLC2701, PLR6301)
  • Reduce return statements in _hf_lang_to_iso_lang_script (7 -> 6)
  • Convert test class methods to plain functions (PLR6301)
  • Import private functions directly from iso_mappings module
  • Remove private names from init.py re-exports
  • format

Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (0fe9fe2)

  • fix: colpali_engine models query processing (#4361)

fix bug (2ea3fe7)

Unknown

  • Update codefuse_models.py (#4363) (1db4399)

  • model: add webAI's ColVec1 Models (#4358)

  • Added webai_models and ViDoRe v3 runner (to be removed) ...

  • Separation between .cache dir and output dir for runpod eval ...

  • Removed run_vdrv3, modified model implementation, TODO: Need to add commit hash after updating model cards ...

  • Added latest commit hash, ready to submit PR ...

  • Apply suggestion from @KennethEnevoldsen

Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>

Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>

Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>

Model meta removal, applied when loading the model?

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>

Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>

Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>

  • Added changes based on pr suggestions ...

  • Fix format for linter check ...

  • Update mteb/models/model_implementations/webai_models.py

  • Update mteb/models/model_implementations/webai_models.py

  • Update mteb/models/model_implementations/webai_models.py

  • fix format


Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (d146a2a)