- Multimodal Large Language Models for Medicine: A Comprehensive Survey [paper]
CVPR'25
- Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation [paper][code]
- Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering [paper]
MedIA'25
- ECAMP: Entity-centered Context-aware Medical Vision Language Pre-training [paper]
TMM'25
- Multi-Grained Vision-and-Language Model for Medical Image and Text Alignment [paper]
NAACL'25
arXiv'25
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features [paper][code]
- MedTrim Meta-Entity Driven Triplet Mining for Aligning Medical Vision-Language Models [paper][code]
- UniBiomed: A Universal Foundation Model for Grounded Biomedical Image Interpretation [paper]
- Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding [paper]
NeurIPS'24
- ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models [paper][code]
- Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning [paper]
arXiv'24