Skip to content

Latest commit

 

History

16 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MCP-MedSAM

Pytorch Implementation of the paper: "MCP-MedSAM: A Powerful Lightweight Medical Segment Anything Model Trained with a Single GPU in Just One Day"

MCP-MedSAM Architecture

📄 Overview

This work proposes a lightweight variant of MedSAM by integrating:

  • A pre-trained Tiny ViT as the vision backbone
  • Two novel prompt types:
    • Modality Prompt
    • Content Prompt
  • A modified mask decoder adapted to these prompts

To further improve performance across imaging modalities, we introduce a modality-aware data sampling strategy that ensures better balance and generalization.

With these enhancements, our model achieves strong multi-modality segmentation performance, and can be trained in approximately 1 day on a single A100 (40GB) GPU.

Requirements

  • Python==3.10.14
  • torch==2.0.0
  • torchvision==0.15.0
  • transformers==4.49.0

Training and Inference

Training and inference can be done by running train.py and infer.py. To note, there is a 'case_data.json' file in the custom dataset class, which is used for speeding up the reading of data, you can replace ? with your local data path. Additionally, we also release the weights of tiny ViT and the whole MCP-MedSAM for inference, which can be downloaded from here. Furthermore, MCP-MedSAM has also been uploaded to the Hugging Face, including pre-trained weights as well.

Citation

@article{lyu2024mcp,
  title={MCP-MedSAM: A Powerful Lightweight Medical Segment Anything Model Trained with a Single GPU in Just One Day},
  author={Lyu, Donghang and Gao, Ruochen and Staring, Marius},
  journal={arXiv preprint arXiv:2412.05888},
  year={2024}
}

About

No description, website, or topics provided.

Resources

Stars

12 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages