2.19.6
2.19.6 (2026-08-20)
Fix
-
fix: skip R_cap division when a query has no relevant docs (#5253)
-
fix: skip R_cap division when a query has no relevant docs
-
test: cover recall_cap with no relevant documents (
ea20326)
Unknown
-
add total parameters to performance size plot (#5245)
-
feat: add total parameters to performance size plot
-
remove test
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (fe08f24)
-
model: add OmniRetriever-7B (#5164)
-
model: add OmniRetriever-7B
-
docs: add OmniRetriever mock-run results
-
fix: address OmniRetriever review feedback
-
Delete mteb_mock_run_results.md
-
format
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (c6801b8)
-
fix: restore the EOS token for e5-mistral (#5237)
Addresses #5230 (in code, we still need the results)
intfloat/e5-mistral-7b-instruct uses last-token pooling — the embedding is the hidden
state of the EOS </s>. Under transformers>=5 that </s> is no longer added, so we
pool the last content token instead and the embeddings are wrong.
The model repo contradicts itself:
tokenizer_config.jsonsets"add_eos_token": truetokenizer.json's post-processor adds only<s>
transformers<5 resolved that in favour of the config and rewrote the post-processor at load
time; transformers>=5 takes tokenizer.json as-is. Nothing warns.
transformers 4.57.6 post_processor: <s> + A + </s>
transformers 5.15.0 post_processor: <s> + A
Passing add_eos_token=True on its own is not a complete fix: on transformers>=5 it
rebuilds the post-processor from add_bos_token, which defaults to False for this repo, so
you get the </s> back but lose the <s>. Both flags are needed.
Results
bBSARDNLRetrieval (nDCG@10), intfloat/e5-mistral-7b-instruct, batch size 8, same GPU.
The before / add_eos_token only / after rows are mteb 2.19.3 with the two tokenizer
kwargs applied through a loader_kwargs override, which is exactly what this patch writes
into the file:
| mteb | transformers | tokenizer | score |
|---|---|---|---|
| 2.18.1 | 4.57.6 | <s> … </s> |
0.37365 |
| 2.18.4 | 4.57.6 | <s> … </s> |
0.37365 |
| 2.18.7 | 4.57.6 | <s> … </s> |
0.37365 |
| 2.19.3 | 4.57.6 | <s> … </s> |
0.37365 |
| before | 5.15.0 | <s> … |
0.06504 |
add_eos_token only |
5.15.0 | … </s> |
0.34506 |
| after | 5.15.0 | <s> … </s> |
0.37360 |
For reference the bBSARD paper reports 0.3770 for this model.
@tomaarsen I will just make you aware of this. I am not sure if it is something that we want to address upstream (if nothing else good to know that it exists).
Re the GritLM → Sentence Transformers migration (#4085) — the GritLM loader scores 0.05508
on main today, it truncates at max_length=512 and it prefixes documents with "Instruct: \nQuery: ". Fixing these give similar scores.
Other models
I scanned to see if there were other cases of this, mostly it found cases of misspecified dependencies or missing. I have fixed those.
The only other case was Salesforce/SFR-Embedding-Code-2B_R, but that doesn't work with never versions of ST and it is unclear
if BeastyZ/e5-R-mistral-7b is intended to run with mean pooling (ships no modules.json or 1_Pooling, so ST falls back to mean pooling)
Co-authored-by: Kenneth <kennethenevoldsen@gmail.com> (a20055d)
-
model: add Gemini Embedding 2 (#5220)
-
model: add Gemini Embedding 2
-
truncate audio inputs to duration limit
-
use query and body fields for text-only retrieval
-
remove tests
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (de40885)