Official implementation of Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation, accepted to CVPR 2026.
Chongyang Xu, Haipeng Li, Shen Cheng, Haoqiang Fan, Ziliang Feng, Shuaicheng Liu†
† Corresponding author.
- 2026-06-08: GAP code, preprocessing scripts, training pipeline, and RoboTwin evaluation wrapper are released.
- 2026-02-27: The paper is available on arXiv.
- 2026-02: GAP is accepted to CVPR 2026.
GAP is a bimanual manipulation policy that reasons over RGB observations with a 3D geometric prior. The policy fuses DINOv3 visual tokens, Pi3 geometry-aware point-map tokens, and robot proprioception, then uses a diffusion decoder to predict both the next action chunk and a future 3D latent target. This action-geometry prediction objective encourages the policy to understand where the scene is going, not only which motor command should come next.
- RGB-only 3D reasoning. GAP uses a pretrained 3D geometric foundation model to obtain geometry-aware latents without requiring explicit depth sensors or point-cloud inputs.
- Joint action and geometry prediction. The policy predicts future actions together with future 3D point-map latents, providing an auxiliary objective aligned with bimanual manipulation.
- RoboTwin-compatible release. The repository keeps RoboTwin as an external dependency and runs evaluation through a local GAP wrapper, so users do not need to modify the official RoboTwin codebase.
- Reproducible scripts. Data preprocessing, training, checkpoint loading, and evaluation are exposed through short shell entry points.
GAP/
├── deploy_policy.py # Policy wrapper used by RoboTwin evaluation
├── deploy_policy.yml # Default evaluation config
├── process_data.sh # Feature extraction entry point
├── train.sh # Training entry point
├── eval.sh # RoboTwin evaluation entry point
├── gap_policy/ # GAP policy, dataset, diffusion modules, configs
├── scripts/
│ ├── process_data.py # HDF5 episodes -> GAP zarr features
│ ├── train.py # Hydra training script
│ └── eval_policy.py # GAP-local copy of RoboTwin eval runner
├── thirdparty/
│ ├── dinov3/ # DINOv3 inference code used by GAP
│ └── pi3/ # Pi3 model code subset used by GAP
└── pretrained/
└── download_weights.sh # Downloads external pretrained weights
The thirdparty/ directory contains the code needed for inference. Large pretrained weights are not tracked by git.
GAP is designed to run inside a working RoboTwin environment. Please first install RoboTwin following its official instructions and verify that the simulator can run a standard policy evaluation.
After activating the RoboTwin environment, clone this repository:
git clone https://github.com/Chongyang-99/GAP.git
cd GAP
export ROBOTWIN_ROOT=/path/to/RoboTwinROBOTWIN_ROOT must point to the complete RoboTwin root directory and contain script/eval_policy.py.
The simulation experiments use RoboTwin's AgileX bimanual embodiment by default, i.e., the ALOHA-AgileX robot model provided in RoboTwin.
GAP expects all pretrained weights under pretrained/:
pretrained/
├── Pi3/
│ ├── config.json
│ └── model.safetensors or pytorch_model.bin
└── dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth
Download them with:
bash pretrained/download_weights.shThe script downloads Pi3 from Hugging Face and the DINOv3 ViT-L checkpoint used by GAP. If Hugging Face access requires authentication, set HF_TOKEN before running the script. The expected weight size is about 4.8 GB in total, so these files are intentionally ignored by git.
Useful overrides:
export GAP_PRETRAINED_ROOT=/path/to/pretrained
export PI3_MODEL_NAME_OR_PATH=/path/to/pretrained/Pi3
export DINOV3_WEIGHTS_PATH=/path/to/dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth
export DINOV3_REPO_DIR=/path/to/GAP/thirdparty/dinov3GAP reads RoboTwin demonstration episodes from:
${ROBOTWIN_ROOT}/data/<task_name>/<task_config>/data/episode*.hdf5
For the demo-clean setting used in our experiments, preprocess one task with:
bash process_data.sh place_dual_shoes demo_clean 100 0Arguments:
process_data.sh <task_name> <task_config> <expert_data_num> <gpu_id>
The script extracts DINOv3 and Pi3 features and saves a zarr dataset to:
data/<task_name>-<task_config>-<expert_data_num>-pi3-20-5.zarr
You can redirect input or output locations without modifying RoboTwin:
export RAW_DATA_ROOT=/path/to/robotwin/data
export OUTPUT_ROOT=/path/to/gap/dataTrain GAP on the preprocessed demonstrations:
bash train.sh place_dual_shoes demo_clean 100 0 0 32 300 100Arguments:
train.sh <task_name> <task_config> <expert_data_num> <seed> <gpu_id> <batch_size> <num_epochs> <checkpoint_every>
Checkpoints are saved to:
checkpoints/<task_name>_<task_config>_<expert_data_num>/<epoch>.ckpt
By default, logging uses offline Weights & Biases mode. To change it:
export WANDB_MODE=onlineEvaluate a trained checkpoint in RoboTwin:
export ROBOTWIN_ROOT=/path/to/RoboTwin
bash eval.sh place_dual_shoes demo_clean demo_clean 100 100 0 "0"Arguments:
eval.sh <task_name> <task_config> <ckpt_setting> <expert_data_num> <checkpoint_num> <gpu_id> <seeds>
The default checkpoint path is:
checkpoints/<task_name>_<ckpt_setting>_<expert_data_num>/<checkpoint_num>.ckpt
To evaluate a specific checkpoint:
export CKPT_PATH=/path/to/checkpoint.ckpt
bash eval.sh place_dual_shoes demo_clean demo_clean 100 100 0 "0"Evaluation results are written inside this GAP repository:
results/<task_name>/GAP/<task_config>/<ckpt_setting>/seed_<seed>/<checkpoint_num>/_result.txt
Set RESULTS_ROOT to use another output directory:
export RESULTS_ROOT=/path/to/resultsThe main training config is:
gap_policy/config/GAP.yaml
Important options:
policy.dinov3_repo_dir: local DINOv3 code path.policy.dinov3_weights_path: DINOv3 checkpoint path.policy.pi3_model_name_or_path: Pi3 checkpoint directory.observation_chunk: temporal observation window, default20.interval: temporal sampling interval, default5.model_3d: 3D backbone name, currentlypi3.
GAP intentionally treats RoboTwin as an external environment:
- preprocessing reads official RoboTwin HDF5 demonstrations;
- evaluation changes into
ROBOTWIN_ROOTso RoboTwin task configs and simulator imports resolve normally; - GAP uses
scripts/eval_policy.pyin this repository to support repository-local result paths; - no files inside the official RoboTwin repository need to be edited.
If you find this repository useful, please cite:
@inproceedings{xu2026gap,
title={Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation},
author={Xu, Chongyang and Li, Haipeng and Cheng, Shen and Fan, Haoqiang and Feng, Ziliang and Liu, Shuaicheng},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year={2026}
}We thank the authors of Pi3, RoboTwin, and Xu et al.'s Diffusion-Based Imaginative Coordination repository for releasing their excellent codebases. GAP builds on RoboTwin for bimanual manipulation evaluation, uses Pi3 for 3D geometric representation, and benefits from Xu et al.'s open-source bimanual policy implementation.