-
Notifications
You must be signed in to change notification settings - Fork 0
Review AnyDexRT
In-Depth Review β AnyDexRT: Calibration-Free Dexterous Hand Retargeting with Few-Shot Human Guidance
Paper: "AnyDexRT: Calibration-Free Dexterous Hand Retargeting with Few-Shot Human Guidance" β arXiv 2607.08341 (Jul 9 2026) Authors: Chenxi Wang, Ying Feng, Hongjie Fang, Shangning Xia, Lixin Yang, Chuan Wen, Cewu Lu Β· Shanghai Jiao Tong University (Cewu Lu group) What it is: a calibration-free, cross-hand humanβrobot retargeting method β the load-bearing L4 bridge of the Dexterous-Hand Data Pyramid, made cheap and hand-agnostic.
Companions: Dexterous-Hand Data Pyramid Β· Do As I Do (the source-data counterpart) Β· One-Hand (cross-hand canonicalization) Β· Dexterous Manipulation.
- Retargeting without per-hand calibration or hand-crafted objectives. AnyDexRT maps human fingertip motion to any dexterous hand by learning task-relevant fingertip correspondences rather than assuming geometric similarity β no precise calibration, no global shape matching.
-
Two-stage mapping: a self-supervised fingertip position mapper (
f_m) + inverse kinematics (f_s), trained with three losses β Partial Chamfer (asymmetric humanβrobot fingertip mapping), Distance Preservation, and Local Motion Preservation (directional consistency in local frames β the key to calibration-robustness). - Few-shot human guidance: operators imitate a tiny set of reference gestures (~5 sampled configs per anchor type, interpolated to ~150β200 paired anchors per hand) to ground the mapping in task-relevant regions.
- Pinch is handled specially: a contact classifier detects human pinch and searches the neighborhood of the mapped robot position for a valid pinch pose β fixing the grasp-critical configs generic mapping misses.
- Results: across 7 dexterous hands, Local Motion Consistency 90.2% (GeoRT 59.8, optimization 52.2) and pinch success 62.0% (GeoRT 29.2); only 3 hyperparameters, runs at 293 Hz, stable under Β±90Β° frame rotations.
- It attacks the pyramid's real bottleneck. The data pyramid argues L4 retargeting fidelity β not raw data volume β is what converts abundant human data into 5-finger skill. AnyDexRT makes L4 calibration-free and cross-hand, so the same pipeline serves Inspire, Allegro, Shadow, LEAP, etc. without per-hand tuning.
- Local-motion preservation is the calibration-robust trick. By enforcing directional consistency in local coordinate frames instead of matching global geometry, it stays stable under large frame-rotation errors (Β±90Β°) β precisely the calibration burden that breaks prior retargeters.
- Complements the source-data side. Pairs naturally with Do As I Do (which produces the human hand-object trajectories) and One-Hand (which canonicalizes hand morphology): AnyDexRT is the cheap universal mapper between them.
flowchart LR
H[Human fingertip motion] --> FM[Fingertip mapper f_m<br/>self-supervised]
subgraph L[losses]
PC[Partial Chamfer<br/>asymmetric humanβrobot]
DP[Distance Preservation]
LMP[Local Motion Preservation<br/>directional, calibration-robust]
end
FM --- L
GUIDE[Few-shot human guidance<br/>~5 configs/anchor β ~150β200 anchors/hand] --> FM
FM --> IK[Inverse kinematics f_s]
CC[Contact classifier<br/>detect pinch β search robot pinch pose] --> IK
IK --> OUT[Robot hand joints Β· 293 Hz]
-
Fingertip mapping (
f_m): learns humanβrobot fingertip correspondence self-supervised. Partial Chamfer maps human fingertip space into the feasible robot fingertip space without requiring full coverage; Distance Preservation keeps pairwise fingertip distances; Local Motion Preservation enforces local-frame directional consistency (less sensitive to calibration than global methods). - Few-shot guidance: operators imitate a small set of reference gestures giving paired human-robot fingertip anchors; an alignment loss minimizes mapped-vs-anchor distance over M pairs. Only two anchor types (lateral rotation, bending); Kβ=5 initial configs interpolated to K=50 (rotation) / K=100 (bending) per finger.
- Contact classifier: binary classifier on finger contact signals; on detected human pinch, searches the neighborhood of the mapped robot position for a valid robotic pinch pose.
-
IK (
f_s): maps target fingertip positions to joint angles.
Motion consistency, averaged across 7 hands (Inspire 6-DoF Β· Ability Β· XHand Β· Wuji Β· Allegro 16-DoF Β· LEAP Β· Shadow 24-DoF):
| Metric | AnyDexRT | GeoRT | Optimization |
|---|---|---|---|
| Global Motion Consistency | 79.9% | 78.3% | 62.0% |
| Local Motion Consistency | 90.2% | 59.8% | 52.2% |
| Pinch success rate | 62.0% | 29.2% | 39.6% |
- Efficiency: 3 hyperparameters; 293 Hz.
- Robustness: stable Local Motion Consistency under Β±90Β° frame rotations; orders-of-magnitude better training stability than GeoRT across initializations.
- Real world: teleoperation on Flexiv Rizon 4 + Wuji Hand, task times e.g. Spray-Bottle 10.6 s, Light-Bulb 17.0 s, Steak-Shoveling 28.0 s, Small-Ball Picking 105.8 s; 8 operators of varying experience.
Significance. AnyDexRT turns the pyramid's L4 bridge from a per-hand, calibration-heavy chore into a cheap, hand-agnostic module β the enabling piece for scaling human data to any 5-finger hardware. Local-motion preservation + contact-aware pinch refinement are the transferable ideas.
Limitations (authors').
- Still needs a few human-guided anchors β future work: automate anchor selection or adapt online from operator feedback.
- Contact refinement is pinch-only β broader contact-rich behaviors need richer contact models.
- Evaluated mainly as teleoperation β training downstream manipulation policies on AnyDexRT-collected data is not yet validated (the true test of its data-collection value).
- Paper: arXiv 2607.08341
- Pyramid placement: L4 (the retargeting bridge) β Dexterous-Hand Data Pyramid
- Counterpart source-data: Do As I Do Β· cross-hand canonicalization: One-Hand Β· glove capture: DexUMI
- Dexterous Manipulation Β· Cross-Embodiment
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)