Skip to content

2.18.3

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 15 Jul 16:56
· 406 commits to main since this release

2.18.3 (2026-07-15)

Fix

  • fix: load gte-Qwen2-7B-instruct in bf16 with trust_remote_code (#4935)

  • fix: load gte-Qwen2-7B-instruct in bf16 with trust_remote_code

Three issues with the stock loader for Alibaba-NLP/gte-Qwen2-7B-instruct:

  1. dtype: the model card recommends bf16; loading in fp16 produces
    numerically-unstable embeddings (nDCG@10 collapses, e.g. ~0.07 vs
    ~0.51 on BRIGHT-Pro biology).
  2. trust_remote_code: the model repo ships modeling_qwen.py with the
    bidirectional attention implementation its embedding head depends
    on. Without trust_remote_code=True sentence-transformers silently
    falls back to stock causal Qwen2 and embeddings collapse to noise.
  3. use_cache: the remote modeling_qwen.py calls
    DynamicCache.get_usable_length(), removed in transformers>=4.56,
    so loading crashes on current transformers once trust_remote_code
    is enabled. Encoding never needs the KV cache;
    config_kwargs={'use_cache': False} skips that code path.
  • Apply suggestions from code review

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (f4a5870)

Unknown

  • model: add StaRSE model (#4936)

  • added starse model meta

  • added starse model meta (ae87966)

  • Set training_datasets for comsat-embed-ja models (empty set + inheritance) (#4930)

Set training_datasets=set() for comsat-embed-ja models

Their own fine-tuning data contains no mteb datasets, so use an empty set
(instead of None) — this lets the base models' training data be inherited
via adapted_from (Qwen/Qwen3-Embedding-8B, cl-nagoya/ruri-v3-310m), giving a
meaningful zero-shot percentage on the leaderboard instead of N/A. (3cdda31)

  • model: Add fusion-embedding-1 (EximiusLabs/fusion-embedding-1-2b-preview) (#4909)

  • model: Add fusion-embedding-1 (EximiusLabs/fusion-embedding-1-2b-preview)

  • Use commit SHA for ModelMeta revision

  • Load via AutoModel trust_remote_code; add fusion-embedding extras group

  • The model repos now ship remote-code files (auto_map in config.json), so the
    wrapper reduces to AutoModel.from_pretrained(..., trust_remote_code=True) and
    no longer needs the model's git package.
  • Add the fusion-embedding optional-dependency group (audio/vision runtime for
    the remote code) and reference it with extra_requirements_groups.
  • Pin revision to the commit carrying the remote code (checkpoint bytes are
    unchanged from the scored tag).
  • training_datasets: list the mteb tasks built on training corpora (AudioCaps,
    FSD50K/FSDKaggle2019, AudioSet) instead of an empty set.
  • n_parameters: 2.8B (frozen base 2.13B + frozen audio tower 0.64B + trained
    connector 16.4M); n_embedding_parameters cleared (the connector is not the
    token-embedding layer); memory_usage_mb recomputed.
  • Point code links at the renamed family repo (Eximius-Labs/fusion-embedding)

  • Add image support to the wrapper

get_image_embeddings routes through the frozen base model's own image path
(exposed by the repo's remote code); images are embedded one at a time, as in
other merged wrappers. modalities now lists image. Fused image+text inputs
raise NotImplementedError.

  • Add video and fused-modality support; define n_embedding_parameters
  • get_video_embeddings: VideoCollator frames through the model's video path
    (the frozen base's own; the repo's remote code exposes embed_video, verified
    bitwise against the raw base model). modalities now lists all four.
  • Fused inputs: element-wise sum of per-modality embeddings, any combination
    (the clip_models.py convention), replacing the NotImplementedError.
  • n_embedding_parameters = 311,164,928 (base token-embedding layer,
    vocab 151,936 x hidden 2,048) — fixes the failing meta test.
  • Revision re-pinned to the commit carrying embed_video; qwen-vl-utils added
    to the fusion-embedding extras (the base's own video preprocessing package).
  • Apply review suggestions; torchcodec-native video path; re-pin revision
  • extra_requirements_groups=["fusion-embedding"] and the meta-field cleanup as
    suggested (max_tokens kept: required ModelMeta field, validation fails
    without it).
  • Video: VideoCollator frame tensors pass directly to embed_video (no image
    splitting); qwen-vl-utils removed everywhere incl. the extras group — the
    remote code reproduces the base's reference preprocessing natively and
    decodes paths with torchcodec's VideoDecoder, verified bitwise-identical to
    both the previous implementation and the raw base model.
  • Revision pinned to the commit carrying embed_video.
  • Exercised end-to-end on MSVDT2VRetrieval (nDCG@10 0.826). (9ea1d63)
  • Add Hanno-Labs/dinghy-law-0.6b-v1 (legal embedding model) (#4926)

  • Add Hanno-Labs/dinghy-law-0.6b-v1 (legal embedding model)

  • Address review: use SentenceTransformerEncoderWrapper + model_prompts (verified byte-identical to q3e loader, 65.85 -> 65.85 on all 8 MTEB(Law) tasks); trim verbose comments to model card

  • Use InstructSentenceTransformerModel with a custom instruction_template (per @Samoed); same mechanism q3e wraps, so 65.85 unchanged; bare per-task instructions via prompts_dict


Co-authored-by: Stephen Solka <stephen@standd.io> (a954de2)

  • InjongoIntent: dropped eng config (#4913)

  • InjongoIntent: dropped eng config (as mteb mirror was missing eng/test split)

  • style: apply ruff format


Co-authored-by: Nicolas Helmeyer <helmeyen@login-4.server.mila.quebec>
Co-authored-by: Nicolas Helmeyer <helmeyen@login-1.server.mila.quebec> (651d53f)

  • Fixup logging message for multimodal sbert models (#4927)

  • fix modalities

  • add log message

  • fix cross-encoder

  • fixup logging message (30ebfee)