-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 Robo3R
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: Manipulation 2 Β· paper #56 Authors: Sizhe Yang, Linning Xu, Hao Li, Juncheng Mu, Jia Zeng, Dahua Lin, Jiangmiao Pang arXiv: 2602.10101 Β· program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1: from RGB frames plus robot state, Robo3R predicts local geometry, relative pose, and a global similarity transformation in a single forward pass (left); its point clouds are visibly cleaner than a depth camera's (middle); and the resulting geometry lifts downstream success rates for imitation learning/sim-to-real, grasp synthesis, and collision-free motion planning (right).
3D input helps manipulation policies, grasp synthesis, and motion planning, but depth cameras (stereo or ToF) are noisy and fail on transparent, reflective, or tiny objects, while general feed-forward reconstruction models (VGGT, ΟΒ³, DepthAnything3, MapAnything) lack manipulation-level geometric precision and reliable metric scale. Robo3R (Shanghai AI Lab / CUHK / USTC / Tsinghua) aims to replace depth sensors and calibration with an RGB-only, manipulation-ready reconstruction model.
Robo3R fuses DINOv2 ViT-L image features (1β2 views) with MLP-encoded robot joint states, processed by 18 alternating global/frame-wise attention blocks. It predicts scale-invariant local point maps via a masked point head (separate robot/object/background branches with depth, ray, and mask heads to avoid over-smoothing), a relative pose head (9D rotation orthogonalized by SVD), and similarity-transformation tokens that map registered points into metric-scale geometry in the canonical robot frame. A keypoint head predicts robot-link keypoint heatmaps; solving PnP against forward-kinematics 3D keypoints refines the camera extrinsics. Training is end-to-end on Robo3R-4M, a new Isaac Sim synthetic dataset of 100,000 scenes / 4 million frames built from 16,911 objects, 4,710 textures, and 6,512 environment maps with extensive domain randomization.
On a held-out synthetic benchmark (2,000 scenes / 80,000 frames), Robo3R reaches monocular point error 0.006 and scale error 0.007 β an order of magnitude below ΟΒ³ (0.061/0.497) and better than MapAnything fine-tuned on the same data (0.010/0.010) β plus relative-pose RTE 0.014 / RRE 0.013 with RTA@0.03 = 0.951. It reconstructs objects as thin as 1.5 mm and handles mirrors and transparent cups that blind a RealSense D455. In real-world downstream tests (Franka Research 3 and bimanual UR5e with XHand, RTX 4090, 10 Hz control): imitation learning with ManiFlow+Robo3R scores 14/16, 15/16, 12/16, 16/16 on Sweep Bean / Insert Screw / Breakfast / BiDex Pour, beating depth-camera, RGB, other feed-forward, and Ο0 baselines; sim-to-real reaches 16/16 (Push Cube) and 12/16 (Pick Cube) vs 7/16 and 5/16 with depth; grasp synthesis and collision-free motion planning gain most on transparent/reflective/small/thin objects (e.g., 5/5 vs 2/5 on thin obstacles).
Positions learned feed-forward reconstruction as a drop-in replacement for depth sensors in manipulation stacks β metric-scale, canonical robot frame, calibration-free β and shows the 3D-representation route to sim-to-real consistency. Related wiki threads: Review-Dexterous-Manipulation Β· Review-VLA-Evaluation.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)