2.16.3
2.16.3 (2026-07-03)
Fix
-
fix(JinaV5OmniWrapper): text-task prefix + channels-last video frames (#4881)
-
fix(JinaV5OmniWrapper): use Query/Document prefix for all text tasks
The text-matching, classification and clustering adapters were trained
WITH the "Query: "/"Document: " prefix, not without it. Running them
without any prefix causes dramatic score drops (e.g. STS12: 0.85→0.38,
Banking77: 0.90→0.30, SprintDuplicateQuestions: 0.96→0.18).
New logic: audio/video non-retrieval tasks skip the prefix (unchanged);
text and image tasks always use the prefix regardless of adapter type.
Update unit tests to reflect the corrected expected behavior.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- fix(JinaV5OmniWrapper): convert torchcodec NCHW video frames to channels-last
torchcodec's VideoDecoder yields (T, C, H, W) uint8 frame batches, but the
model's HF remote code detects video only for channels-last (T, H, W, 3|4)
arrays. The mismatched tensor fell through to the text fallback and was
embedded as str(array) — every video collapsed to near-identical garbage
embeddings (e.g. Shot2Story20KAT2VRetrieval nDCG 0.0009).
Verified on GPU: same video as channels-first tensor vs channels-last array
gives cosine 0.14 between embeddings; the channels-first embedding matches
the embedding of the array's string repr.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- lint
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (598a7eb)
Unknown
fix(ebind): pad sub-window audio clips so short samples don't crash the batch (3fdcb0d)
-
model: Add Byrne-Embed model implementation (#4847)
-
Add Byrne-Embed model implementation (Quazim0t0/Byrne-Embed)
-
Byrne-Embed: load via trust_remote_code AutoModel (drop snapshot_download)
-
Update mteb/models/model_implementations/byrne_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: SeanceTable <apoetyouknow@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (8b8f169)