Skip to content

2.12.25

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 20 Apr 11:38
· 741 commits to main since this release

2.12.25 (2026-04-20)

Fix

  • fix: drop unused modality columns in dataloader for cross-modal tasks (#4440)

  • fix: handle None text/image in multimodal retrieval tasks

Cross-modal retrieval tasks (CIRRIT2IRetrieval, NIGHTSI2IRetrieval,
Fashion200kI2TRetrieval, VisualNewsI2TRetrieval) have corpus/query
items where text or image can be None for single-modality entries.

  • _corpus_to_dict: handle None text/title gracefully
  • _custom_collate_fn: allow None values in batches instead of raising
  • _combine_queries_with_instruction_text: skip string ops on None text
  • random_baseline: skip None items when encoding each modality

Closes #4436

  • fix: handle None query text in dataloader instead of downstream models

Normalize None text to "" in _combine_queries_with_instruction_text,
matching the existing pattern in _corpus_to_dict. Revert random_baseline
and collation changes as they're no longer needed.

  • fix: drop unused modality columns in dataloader to prevent None errors

Cross-modal retrieval tasks have None values for modalities not used by
that side of the retrieval (e.g. text=None in image-only corpus for it2i
tasks). Instead of adding None-guards throughout the collate function and
models, drop columns for modalities not needed for the current prompt
type in _prepare_dataset. The task category (e.g. it2i) already encodes
which modalities each side needs.

Closes #4436

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>


Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> (eaf2c9d)