The official repository of our paper "MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding".
You can explore our demo by running demo.ipynb. This demonstration illustrates how our MA-LMM serves as a plug-and-play module that can be integrated into InstructBLIP seamlessly, requiring no fine-tuning for zero-shot evaluation.
You can install the conda environment by running:
git clone https://github.com/boheumd/MA-LMM.git
cd MA-LMM
pip install -e .If you are running the code on Apple Silicon, you need to use eva-decord instead of decord. Here is the modification in the requirements.txt file you should do:
contexttimer
eva-decord
einops>=0.4.1
fairscale==0.4.4
...
Before running pip install -e ., ensure you have the correct requirements.
For the long-term video understanding task, we conduct experiments including (LVU) and two standard video summarization datasets (Breakfast, COIN).
For the video question answering task, we conduct experiments including MSRVTT, MSVD, and ActivityNet. For the video captioning task, we also conduct experiments on Youcook2 dataset.
You can download videos for each dataset through the script provided here (lavis/datasets/download_scripts). For LVU/Breakfast/COIN datasets, please download the original videos through the official link provided above.
Then extract video frames of each video with fps=10. Example preprocess code is provided here extract_frames.py. Since different FFMPEG versions are used, the actual extracted frame lengths can be slightly inconsistent. You may need to update the actual frame_length for each video in the annotation file.
├── data
└── activitynet
├── annotation
├── frames
├── videos
└── breakfast
├── annotation
├── frames
├── videos
└── coin
├── annotation
├── frames
├── videos
└── lvu
├── annotation
├── frames
├── videos
└── msrvtt
├── annotation
├── frames
├── videos
└── msvd
├── annotation
├── frames
├── videos
└── youcook2
├── annotation
├── frames
├── videos
We use Vicuna-v1.1 as our pre-trained LLM weights, you can download from this link as arrange in this format.
├── llm
├── vicuna-7b
├── vicuna-13b
Our model leverages pre-trained weights from InstructBlip, which was only pre-trained on image-text pairs. Our training process occurred on four A100 GPUs. If you would like to fine-tune the model for various video datasets, please run the following command:
bash run_scripts/${dataset}/train.sh # Please choose the task from the following list
# ['director', 'genre', 'relationship', 'scene', 'way_speaking', 'writer', 'year']
datasets.lvu_cls.task ${task}We also provided finetuned checkpoints for each video dataset. Please download the saved_model.tar and unzip it. For the test script corresponding to each dataset, provide the path to the extracted checkpoint to execute the evaluation.
bash run_scripts/${dataset}/test.sh ${checkpoint_path}Our model can also leverage pre-trained weights from InstructBlip without any finetuning to conduct zero-shot evaluation on video datasets.
bash run_scripts/${dataset}/test.shOne important hyper-parameters memory_bank_length, please change that in the training script on different datasets.
# pre-defined length of the memory bank
model.memory_bank_length ${value}
# value=0 means without using the memory bankThe core algorithm for the memory bank compression algorithm is here.
If you find our code or our paper useful for your research, please [★star] this repo and [cite] the following paper:
@inproceedings{he2024malmm,
title = {MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding},
author = {He, Bo and Li, Hengduo and Jang, Young Kyun and Jia, Menglin and Cao, Xuefei and Shah, Ashish and Shrivastava, Abhinav and Lim, Ser-Nam},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2024}
}We referenced the repo below for the code
본 프로젝트는 Vision, Audio 등 다양한 모달리티를 하나의 임베딩 공간에 결합하는 ImageBind 모델을 기반으로, 장기 기억 및 추론 능력을 갖춘 멀티모달 기억 시스템을 개발하는 것을 목표로 합니다.
기존 대형언어모델(LLM)의 한계인 장기 기억 부족과 추론 오류를 극복하기 위해, **MA-LMM (Memory-Augmented Large Multimodal Model)**의 아키텍처를 차용하고, 여기에 ImageBind의 강력한 멀티모달 인코딩 능력을 결합했습니다.
본 프로젝트에서 구현된 모델은 다음과 같은 독자적인 파이프라인 아키텍처를 가집니다.
- Image Encoder: ImageBind-Huge의 Vision Encoder (출력: 1024차원)
- Text Encoder: ImageBind-Huge의 Text Encoder (출력: 1024차원)
- FC Layer (imagebind_fc): ImageBind의 Vision, Text 임베딩을 결합한 2048차원 벡터를 받아, Q-Former가 처리할 수 있는 1024차원 벡터로 변환하는
nn.Linear(2048, 1024)레이어 - 순서 임베딩 (turn_pe): Visual Dialog의 턴(turn) 순서 정보(0~10)를 학습하기 위한
nn.Embedding(11, 1024)레이어
- Q-Former: Vision 특징과 언어(질문)를 연결하는 핵심 모듈. 입력으로 1024차원의 특징 벡터를 받도록 hidden_size=1024로 설정
- 언어 생성 모듈 (Frozen)
- Projection Layer (llm_proj): Q-Former의 출력(1024차원)을 LLM의 입력 차원(2048차원)으로 변환하는
nn.Linear(1024, 2048)레이어 - LLM: Llama-3.2-1B 모델을 bitsandbytes를 통해 4-bit 양자화하여 로드
- Projection Layer (llm_proj): Q-Former의 출력(1024차원)을 LLM의 입력 차원(2048차원)으로 변환하는
ImageBind와 MA-LMM의 라이브러리 의존성 충돌(PyTorch, timm 버전 등)을 해결하기 위해, 각자의 역할을 독립된 가상환경에서 수행하는 2단계 파이프라인 방식을 채택했습니다.
ImageBind 전용 가상환경에서, VisDial v1.0 데이터셋의 모든 이미지와 각 대화 턴의 질문 텍스트에 대한 Vision/Text 임베딩을 미리 추출하여 디스크에 .pt 파일로 저장합니다.
MA-LMM 학습용 가상환경에서, 사전 추출된 임베딩을 입력으로 받아 오직 imagebind_fc와 turn_pe 레이어만 학습시킵니다.
conda env create -f environment_imagebind.yml
conda activate imagebind_envconda env create -f environment_malmm.yml
conda activate malmm참고: environment_*.yml 파일은 conda env export > [파일이름].yml 명령어로 생성할 수 있습니다.
• 데이터셋 다운로드 lavis/datasets/visdial/ 폴더에 VisDial v1.0 어노테이션과 MS COCO 2014, VisualDialog_val2018 이미지를 다운로드한 후 압축 해제 • 임베딩 추출
conda activate imagebind_env
python extract_visdial_embeddings.py• 모델 파인튜닝
conda activate malmm
cd MA-LMM
python -m lavis.tasks.run –cfg-path lavis/configs/tasks/finetune_visdial.yaml• 성능 평가
python evaluate.py \
–checkpoint “lavis/output/finetune_visdial/finetune_visdial/checkpoint_latest.pth” \
–llm_model_path “llm/llama-3.2-1B/Llama-3.2-1B”-- checkpoint 에, 성능을 평가하고자 체크포인트 경로를 적절히 지정합니다. ! 주의할 점은, cc3m dataset의 경우 lavis/models/../blip2.py의, 208 line의
encoder_config.hidden_size = 1024 # 768 -> 파인튜닝 이전 결과를 위해서, 768으로 돌려야 함.모델 BLEU-4 Score 파인튜닝 전 (FC Layer만 학습) 14.96 파인튜닝 후 (VisDial로 FC+PE 학습) 24.80
분석: VisDial 데이터셋과 순서 임베딩을 이용한 파인튜닝을 통해 BLEU-4 점수가 약 65.8% 향상되었습니다.
• Retrieval 로직 구현: Recency, Frequency, Saliency 기반 스코어링 함수를 구현하여 대화 맥락에 맞는 과거 기억을 동적으로 검색하는 모듈 추가
• Self-Reflection 적용: 검색된 기억을 <reminder> 태그로 프롬프트에 주입하고, 모델이 연쇄적 사고(CoT)를 통해 더 논리적인 답변 생성
• 종합 평가: VisDial뿐만 아니라 ScienceQA 등 다양한 벤치마크 데이터셋으로 정량·정성 평가 수행

