Skip to content

3. Vision‐Language Models for Image Classification

athenas-lab edited this page Apr 5, 2025 · 8 revisions

Vision-Language Models

CLIP

  • A multimodal model that learns the correlation between image and text using contrastive learning.

ViLBERT

Clone this wiki locally