Skip to content

2.10.7

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 06 Mar 16:02
· 917 commits to main since this release

2.10.7 (2026-03-06)

Fix

  • fix: Code leaderboard is failing (#4207)

fix code leaderboard (dec66d6)

Unknown

  • model: Add nomic-ai/nomic-embed-multimodal-7b dense embedding model (#4186)

  • feat: Add nomic-ai/nomic-embed-multimodal-7b dense embedding model

Add BiQwen2_5Wrapper and ModelMeta for nomic-embed-multimodal-7b,
a dense (single-vector) multimodal embedding model for visual
document retrieval using cosine similarity.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

  • fix: Use correct set type for TRAINING_DATA

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

  • fix: Add nomic-embed-multimodal-7b to _MISSING_N_EMBEDDING_MODELS

PEFT adapter repo has no config.json or model.safetensors,
so _from_hub cannot extract n_embedding_parameters.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

  • fix: Use processor.score instead of custom similarity in BiQwen2_5Wrapper

Remove custom similarity method from BiQwen2_5Wrapper to use the
processor's built-in scoring functionality, following the established
pattern used by other ColPali models.

Addresses review feedback in PR #4186.

Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>

  • fix: Use parent class methods instead of custom overrides in BiQwen2_5Wrapper

Remove custom get_image_embeddings and get_text_embeddings methods
to use the inherited ColPaliEngineWrapper implementations. For dense
embedding models with fixed-size vectors, the parent class methods
(extend + pad_sequence) produce equivalent results to the custom
implementation (append + torch.cat).

Also remove unused tqdm import.

Addresses second review feedback in PR #4186.

Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>

  • feat: Add ModelMeta for nomic-ai/nomic-embed-multimodal-3b
  • Add nomic_embed_multimodal_3b ModelMeta with 3B parameters
  • Update BiQwen2_5Wrapper default to use 3B model for better accessibility
  • Scale memory usage proportionally (6200 MB vs 14400 MB for 7B)
  • Maintain same architectural specs (embed_dim=128, max_tokens=128000)

Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>

  • feat: Add nomic-ai/nomic-embed-multimodal-3b to test exceptions
  • Add nomic-embed-multimodal-3b to _MISSING_N_EMBEDDING_MODELS list
  • Update test configuration to handle PEFT adapter repo structure
  • Matches pattern of existing 7B model exception

Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>

  • fix nomic

  • add 3b revision and lint

  • add fixes

  • rem

  • cleanup

  • Update mteb/models/model_implementations/nomic_multimodal.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • add embedding parameters

  • add base revision

  • remove exceptions from test

  • added training data

  • lint

  • Update mteb/models/model_implementations/nomic_multimodal.py

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>


Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Your Name <you@example.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (b398ea7)

  • dataset: Update vidore tasks to include the OCR'd text (#4191)

  • dataset: Add OCR adaption of vidore tasks

  • Add beta tag to the new tasks

  • add nuclear and telecom

  • Apply suggestions from code review

  • Update mteb/tasks/retrieval/multilingual/vidore3_bench_retrieval.py

  • updated description and added version

  • add superseeded from

  • fix imports

  • fix init

  • fix private test

  • add some tasks statics

  • add nuclear


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (9a6b98f)

  • final reupload tasks from mmteb (#4200)

final reupload from mmteb (7037bfe)