Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

17 Commits
 
 

Repository files navigation

Awesone-Vision-Language-Models

Table of Contents

Survey

  • Multimodal Large Language Models for Medicine: A Comprehensive Survey [paper]

Papers

2025

CVPR'25

  • Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation [paper][code]
  • Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering [paper]

MedIA'25

  • ECAMP: Entity-centered Context-aware Medical Vision Language Pre-training [paper]

TMM'25

  • Multi-Grained Vision-and-Language Model for Medical Image and Text Alignment [paper]

NAACL'25

  • VividMed: Vision Language Model with Versatile Visual Grounding for Medicine [paper][code]

arXiv'25

  • SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features [paper][code]
  • MedTrim Meta-Entity Driven Triplet Mining for Aligning Medical Vision-Language Models [paper][code]
  • UniBiomed: A Universal Foundation Model for Grounded Biomedical Image Interpretation [paper]
  • Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding [paper]

2024

NeurIPS'24

  • ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models [paper][code]
  • Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning [paper]

arXiv'24

  • POINTS: IMPROVING YOUR VISION-LANGUAGE MODEL WITH AFFORDABLE STRATEGIES [paper][code]
  • CONTRASTIVE LOCALIZED LANGUAGE-IMAGE PRE-TRAINING [paper]
  • ARIA : An Open Multimodal Native Mixture-of-Experts Model [paper][code]

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors