Skip to content

ruihangzhang97/proxypose

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

1 Commit
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🎯 ProxyPose

ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

arXiv Project Page

Ruihang Zhang*1, Felix Taubner*1,2, Pooja Ravi1, Kiriakos N. Kutulakos1,2, David B. Lindell1,2

1University of Toronto   2Vector Institute   *Equal contribution


TL;DR: One query pixel in, a full 6‑DoF pose trajectory out.


⚡️ Quick start

🛠️ 1. Install

# Clone the repository
git clone https://github.com/ruihangzhang97/proxypose.git
cd proxypose

# Create and activate a conda environment
conda create -n proxypose python=3.10 -y
conda activate proxypose

# Install PyTorch 
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121

# Install PyTorch3D 
export FORCE_CUDA=1
pip install --no-build-isolation git+https://github.com/facebookresearch/pytorch3d.git@stable

# Install ProxyPose and all remaining dependencies
pip install -e .

📦 2. Download model weights

The weights are downloaded automatically on first run — no manual steps needed.

Weight HuggingFace Size
Wan2.1-T2V-14B (base) Wan-AI/Wan2.1-T2V-14B ~30 GB
ProxyPose LoRA ruihangzhang79/proxypose ~600 MB

🖱️ 3. Pick your prompt point

proxypose-annotate --input-video video/my_video.mp4

Opens http://localhost:7860 in your browser. Click on a query point in the first frame, press Save. Coordinates are written to video/my_video.points.json.

🔭 4. (Optional) Estimate focal length with Depth Anything 3

By default, we assume a 45° horizontal field of view. For improved accuracy, consider using Depth Anything 3.

First, install Depth Anything 3 inside your ProxyPose environment:

git clone https://github.com/ByteDance-Seed/Depth-Anything-3.git
cd Depth-Anything-3

pip install --no-build-isolation -e .

Then, from the ProxyPose repository, run:

# Requires:  pip install hatchling editables
python -m inference.annotation.depth_anything video/my_video.mp4

✅ 5. Run inference

Pass the query JSON saved by the annotator directly to --prompt:

proxypose-infer \
    --video_path   video/my_video.mp4 \
    --output_path  output/result.mp4 \
    --prompt       video/my_video.points.json \
    --depth_anything_path video/my_video.da3.npz   # optional, omit to use fixed 45° FOV

📖 Citation

@article{zhang2026proxypose,
  title={ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation},
  author={Ruihang Zhang and Felix Taubner and Pooja Ravi and Kiriakos N. Kutulakos and David B. Lindell},
  journal={arXiv preprint arXiv:2607.06555},
  year={2026}
}

About

Official repository for the paper "ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation"

Resources

Stars

199 stars

Watchers

3 watching

Forks

Releases

No releases published

Packages

 
 
 

Contributors