Junqi You1, Chieh Hubert Lin2, Weijie Lyu2, Zhengbo Zhang3, Ming-Hsuan Yang2
1Shanghai Jiao Tong University, 2UC Merced, 3Singapore University of Technology and Design
We propose InstaInpaint, a reference-based feed-forward framework that produces 3D-scene inpainting from a 2D inpainting proposal within 0.4 seconds. InstaInpaint achieves a 1000× speed-up from prior methods while maintaining a state-of-the-art performance across two standard benchmarks.
# 1. Clone the repository
git clone https://github.com/dhmbb2/InstaInpaint.git
cd InstaInpaint
# 2. Create conda environment
export CUDA_HOME=/usr/local/cuda-12.4/
conda create -n instainpaint python=3.10
conda activate instainpaint
# 3. Install PyTorch
pip install torch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 --index-url https://download.pytorch.org/whl/cu124
# 4. Install dependencies
pip install -r requirements.txtNote: If you are using a different CUDA version, update the pre-built flash-attn wheel in requirements.txt accordingly.
Download the preprocessed Spin-NeRF dataset and checkpoint from Hugging Face.
mkdir -p data checkpoints
wget -O data/spinnerf_dataset.zip https://huggingface.co/dhmbb/instainpaint/resolve/main/spinnerf_dataset.zip?download=true
unzip -q data/spinnerf_dataset.zip -d data/
wget -O checkpoints/exp_ins+random+3d_121_multimask.pth https://huggingface.co/dhmbb/instainpaint/resolve/main/exp_ins%2Brandom%2B3d_121_multimask.pth?download=trueCUDA_VISIBLE_DEVICES=0 bash scripts/test_inpaint_spinnerf.shThe inference results are written to outputs/spinnerf_eval.
Download DL3DV_960P from the official repository and put it in data/dl3dv_960.
mkdir -p checkpoints
wget -O checkpoints/pretraining_ckpt.pth https://huggingface.co/dhmbb/instainpaint/resolve/main/pretraining_ckpt.pth?download=trueNext, we need to use SAM2 to get instance masks and cache them for training efficiency.
# create another environment for sam2
conda create -n sam python=3.10
conda activate sam
cd third-party/sam2
pip install -e .
cd ../../
wget -O checkpoints/sam2.1_hiera_large.pt https://dl.fbaipublicfiles.com/segment_anything_2/092824/sam2.1_hiera_large.pt
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 bash scripts/precalculate_mask.sh We also need to cache image depth from the pretrained LRM in order to calculate regional masks during training.
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 bash scripts/precalculate_depth.shCUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 WORLD_SIZE=8 bash scripts/train_inpaint_4v_clip15.sh instainpaint_trainIf you find this work useful in your research, please consider citing:
@article{you2025instainpaint,
title={Instainpaint: Instant 3d-scene inpainting with masked large reconstruction model},
author={You, Junqi and Lin, Chieh Hubert and Lyu, Weijie and Zhang, Zhengbo and Yang, Ming-Hsuan},
journal={arXiv preprint arXiv:2506.10980},
year={2025}
}