Skip to content

RSS 2026 SimToolReal

hwoo.han edited this page Aug 9, 2026 · 2 revisions

SimToolReal: An Object-Centric Policy for Zero-Shot Dexterous Tool Manipulation

Venue: RSS 2026 (Sydney, Jul 13–17) Β· Session: RL Β· paper #151 Authors: Kushal Kedia, Tyler Ga Wei Lum, Jeannette Bohg, Karen Liu arXiv: 2602.16863 Β· program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

SimToolReal zero-shot tool use (Figure 1 of arXiv 2602.16863, Β© the authors)

Top: a single policy deployed zero-shot on novel real tools and tasks β€” hammer, marker, spatula over a pan, and brushes β€” on a KUKA arm with a five-fingered hand. Bottom: the characteristic grasp β†’ in-hand rotation β†’ tool-use sequence for sweeping crumpled paper into a dustpan with a brush.

Problem

Tool use is a hard class of dexterity β€” grasping thin objects lying flat, in-hand reorientation into functional poses, and forceful contact β€” that parallel-jaw grippers resist and teleoperation captures poorly. Prior sim-to-real RL needed per-task object modeling and reward tuning, limiting generality.

Method

SimToolReal trains a single goal-conditioned, object-centric RL policy in simulation over procedurally generated tool-like primitives, with the universal objective of moving each object through random goal poses (poses represented via D = 4 keypoints); this induces grasping, in-hand rotation, and stable-contact skills without task-specific engineering. Training uses SAPG optimization with an asymmetric critic. At deployment, SAM 3D recovers the object mesh and graspable-region bounding box from RGB-D, and goal-pose trajectories are extracted from a human video, so the sim-trained policy runs zero-shot on real tools. Hardware: 22-DoF Sharpa five-fingered hand on a 7-DoF KUKA iiwa 14 (29-DoF control).

Results

On the authors' DexToolBench (6 tool categories, 12 instances, 24 task trajectories, 120 real rollouts at 5 trials each), the single policy shows strong zero-shot Task Progress; on brush-sweeping variants it scores 98.0% (no rotation) and 82.7% (with 90Β° rotation) versus 61.0/10.8% for Fixed Grasp and 8.1/0% for Kinematic Retargeting β€” the claimed 37% average improvement. In simulation it matches specialist per-task RL policies on their own training setups while specialists collapse under object or trajectory changes. Failure modes: pose-tracking loss 43.7%, drops 34.5%, incomplete in-hand rotation 18.2%, grasp failure 3.6%, with consistent re-grasp recovery behavior.

Significance

Shows that "reach random goal poses with random primitives" is a sufficient universal pretext task for dexterous tool use, replacing per-task sim-to-real engineering with one policy steered at test time by human-video trajectories. Related wiki threads: RL Β· Review-Dexterous-Manipulation.

← Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally