Repository navigation
2.21.9
2.21.9 (2026-09-29)
Documentation
Fix
-
fix: repair NanoVDR document encoding path (#5552)
-
fix: repair NanoVDR document encoding path
nanovdr/NanoVDR-S-Multi fails at document-encoding time with an
ImportError, and has done since beee210 (#4699). Two bugs, both past
_load_teacher():
-
_load_teacherimports_build_qwen3_vl_for_embedding_classfrom
qwen3_vl_embedding_models, but #4699 deleted that factory when it
movedQwen3VLEmbeddingWrapperfromAbsEncoderto
InstructSentenceTransformerModel. Recovered it verbatim from
beee2102^and vendored it here: it was private, and this wrapper is
its only remaining consumer, so re-adding it to that module would
partly undo #4699. Only additions are the type annotations current
lint requires;Cache/Qwen3VLConfigstay under TYPE_CHECKING so
the module remains importable without torch (#5463). -
_encode_queriespassedconvert_to_numpy=Falsewhile leaving
convert_to_tensorat its default, which sentence-transformers
documents as returninglist[Tensor].mteb._convert_to_tensorthen
raises "only one element tensors can be converted to Python scalars".
Nowconvert_to_tensor=True, plus.cpu()to match
_encode_documents--cos_simdoes no device harmonisation, so a
CUDA query matrix against a CPU corpus matrix would fail.
The import is lazy and only runs when encoding documents, so
mteb.get_model() and query encoding both worked and the breakage went
unnoticed for ~3.5 months.
Verified end to end on VidoreTabfquadRetrieval: the teacher loads, the
corpus encodes, and the run completes past the similarity step that
previously raised.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- refactor: drop the vendored Qwen3VL wrapper
The vendored _build_qwen3_vl_for_embedding_class turns out to be
redundant. Qwen3VLModel is already the LM-head-less trunk and its
forward returns Qwen3VLModelOutputWithPast.last_hidden_state, which is
all the pooling in _encode_documents needs. Its
_checkpoint_conversion_mapping = {} override was dead weight too: that
attribute is not set on Qwen3VLPreTrainedModel, Qwen3VLModel, or
Qwen3VLForConditionalGeneration in current transformers.
The one thing the wrapper did provide was resetting rope_deltas per
forward -- Qwen3VLModel.forward reads self.rope_deltas but never
resets it -- so that moves to the call site.
Verified against the previous commit:
from_pretrainedloads 625/625 weights with no init warnings, so the
checkpoint keys matchQwen3VLModeldirectly- same input gives a bitwise identical
last_hidden_state(max abs
diff 0.0) MockAny2AnyRetrievalT2Igives identical values for all 149 metrics
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (75b4800)
Refactor
- refactor: move DS1000Retrieval to v2 dataset format (#5542)
Point DS1000Retrieval at the re-uploaded mteb/DS1000Retrieval dataset
and drop the custom load_data, same as #5495.
Unknown
-
fix bugs in docs (#5535)
-
fix bugs in docs
-
merged other task types
-
fix links related to get_tasks (
d98f640) -
Revert "dataset: add MM-BRIGHT retrieval tasks" (#5531)
Revert "dataset: add MM-BRIGHT retrieval tasks (#5168)"