2.12.20
2.12.20 (2026-04-16)
Fix
-
fix: handle Transformers v5 BaseModelOutputWithPooling return types i… (#4328)
-
fix: handle Transformers v5 BaseModelOutputWithPooling return types
Transformers v5 changed get_text_features, get_image_features, and
get_audio_features to return BaseModelOutputWithPooling instead of
plain tensors. This caused AttributeError when tensor operations
like .norm() were applied directly to the output.
Added isinstance(output, BaseModelOutputWithPooling) checks to
extract pooler_output when needed, maintaining backward compatibility
with Transformers v4 tensor returns.
Affected model wrappers:
- clap_models.py: text path (audio path already handled)
- align_models.py: text and image paths
- wav2clip_model.py: text path (CLIP encoder)
- llm2clip_models.py: text and image paths
- siglip_models.py: text and image paths (previously accessed
.pooler_output directly without fallback)
Closes #4081
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- lint
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (367d554)
- fix: double retrieval dataset loading (#4399)
fix retrieval dataset loading (316fca3)