Skip to content

2.19.6

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 20 Aug 16:22
· 236 commits to main since this release

2.19.6 (2026-08-20)

Fix

  • fix: skip R_cap division when a query has no relevant docs (#5253)

  • fix: skip R_cap division when a query has no relevant docs

  • test: cover recall_cap with no relevant documents (ea20326)

Unknown

  • add total parameters to performance size plot (#5245)

  • feat: add total parameters to performance size plot

  • remove test


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (fe08f24)

  • model: add OmniRetriever-7B (#5164)

  • model: add OmniRetriever-7B

  • docs: add OmniRetriever mock-run results

  • fix: address OmniRetriever review feedback

  • Delete mteb_mock_run_results.md

  • format


Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (c6801b8)

  • fix: restore the EOS token for e5-mistral (#5237)

Addresses #5230 (in code, we still need the results)

intfloat/e5-mistral-7b-instruct uses last-token pooling — the embedding is the hidden
state of the EOS &lt;/s&gt;. Under transformers&gt;=5 that &lt;/s&gt; is no longer added, so we
pool the last content token instead and the embeddings are wrong.

The model repo contradicts itself:

  • tokenizer_config.json sets &#34;add_eos_token&#34;: true
  • tokenizer.json's post-processor adds only &lt;s&gt;

transformers&lt;5 resolved that in favour of the config and rewrote the post-processor at load
time; transformers&gt;=5 takes tokenizer.json as-is. Nothing warns.

transformers 4.57.6  post_processor: &lt;s&gt; + A + &lt;/s&gt;
transformers 5.15.0  post_processor: &lt;s&gt; + A

Passing add_eos_token=True on its own is not a complete fix: on transformers&gt;=5 it
rebuilds the post-processor from add_bos_token, which defaults to False for this repo, so
you get the &lt;/s&gt; back but lose the &lt;s&gt;. Both flags are needed.

Results

bBSARDNLRetrieval (nDCG@10), intfloat/e5-mistral-7b-instruct, batch size 8, same GPU.
The before / add_eos_token only / after rows are mteb 2.19.3 with the two tokenizer
kwargs applied through a loader_kwargs override, which is exactly what this patch writes
into the file:

mteb transformers tokenizer score
2.18.1 4.57.6 &lt;s&gt; … &lt;/s&gt; 0.37365
2.18.4 4.57.6 &lt;s&gt; … &lt;/s&gt; 0.37365
2.18.7 4.57.6 &lt;s&gt; … &lt;/s&gt; 0.37365
2.19.3 4.57.6 &lt;s&gt; … &lt;/s&gt; 0.37365
before 5.15.0 &lt;s&gt; … 0.06504
add_eos_token only 5.15.0 … &lt;/s&gt; 0.34506
after 5.15.0 &lt;s&gt; … &lt;/s&gt; 0.37360

For reference the bBSARD paper reports 0.3770 for this model.

@tomaarsen I will just make you aware of this. I am not sure if it is something that we want to address upstream (if nothing else good to know that it exists).

Re the GritLM → Sentence Transformers migration (#4085) — the GritLM loader scores 0.05508
on main today, it truncates at max_length=512 and it prefixes documents with &#34;Instruct: \nQuery: &#34;. Fixing these give similar scores.

Other models

I scanned to see if there were other cases of this, mostly it found cases of misspecified dependencies or missing. I have fixed those.
The only other case was Salesforce/SFR-Embedding-Code-2B_R, but that doesn't work with never versions of ST and it is unclear
if BeastyZ/e5-R-mistral-7b is intended to run with mean pooling (ships no modules.json or 1_Pooling, so ST falls back to mean pooling)

Co-authored-by: Kenneth <kennethenevoldsen@gmail.com> (a20055d)

  • model: add Gemini Embedding 2 (#5220)

  • model: add Gemini Embedding 2

  • truncate audio inputs to duration limit

  • use query and body fields for text-only retrieval

  • remove tests


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (de40885)