-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 GHOST
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: Manipulation 2 Β· paper #55 Authors: Sriram Krishna, Ben Eisner, Haotian Zhan, Ying Yuan, Haoyu Zhen, Chuang Gan, Shubham Tulsiani, David Held arXiv: 2606.10025 Β· program page
Summary compiled from the arXiv paper (v1, CMU Robotics Institute + UMass Amherst); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1: robot teleop demos on multiple tasks (top) train both the high-level 3D goal predictor Ο_hi and the low-level goal-conditioned policy Ο_lo, while optional human demos on a novel task (left) train only Ο_hi; at rollout on a novel task (bottom), the system alternates Plan (Ο_hi predicts a 3D end-effector sub-goal, visualized on the point cloud) and Act (Ο_lo executes toward it).
Imitation-learned visuomotor policies struggle to generalize beyond the training distribution, and incorporating human video normally requires noisy action retargeting into the robot's embodiment. The paper asks whether factorizing control into an embodiment-agnostic sub-goal predictor and an embodiment-specific controller improves both in-distribution performance and out-of-distribution transfer.
GHOST splits the policy into (i) a high-level Ο_hi β a decoder-only transformer over frozen DINOv3 patch tokens from multi-view RGB-D (patch tokens augmented with 3D coordinates, plus gripper-pose, Flan-T5 language, and a learnable human/robot embodiment token) that predicts a dense per-patch GMM over the next sub-goal, represented as 4 end-effector 3D keypoints (gripper base, two fingertips, grasp center; MANO palm/thumb/index/grasp-center for human hands) β and (ii) a low-level Ο_lo, a Diffusion Policy conditioned on a simple spatial interface: predicted 3D keypoints projected into each camera and converted to dense distance-field end-effector heatmaps. Sub-goals are extracted automatically at gripper open/close transitions in robot teleop data (manually annotated for human demos); human demonstrations (hand poses from off-the-shelf trackers, scale-resolved with Grounded-SAM + depth) train only Ο_hi, keeping Ο_lo purely robot-trained. Tasks use 17β50 demos each; evaluation is 30 trials per method with bootstrap CIs.
Hierarchy alone helps in-distribution: on fold-onesie, final success jumps from 10% (flat Diffusion Policy) to 80% (GHOST robot-only), and on hammer-pin from 16.7% to 50%. Human demos then unlock OOD transfer: 63.3% on mug-on-table (novel object combination, vs 13.3% DP / 28.3% MimicPlay), 56.7% final on fold-onesie-ood (novel instance, vs 0% MimicPlay), 36.7% on fold-towel (novel category + skill composition, vs 16.7% MimicPlay), and 70% on hammer-pin with a novel target pin. An oracle ablation shows Ο_lo generalizes zero-shot to towels (40%β90% when Ο_hi gets robot-quality sub-goals), locating the OOD bottleneck in the high-level human-robot visual domain gap.
A clean demonstration that 3D sub-goals are a practical embodiment-agnostic interface for folding human video into robot skill learning without action retargeting β directly relevant to the human-data threads in Review-Human-Video-Transfer and the hierarchical-policy discussions in Review-Dexterous-Manipulation and Review-VLA-Evaluation.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)