-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 DexNDM
Venue: ICLR 2026 Category: Dexterous Manipulation Trend tag: Trend 6
flowchart LR
Exp[Category RL experts<br/>privileged obs, PPO] --> Gen[Distill to single<br/>generalist policy BC]
Real[Autonomous real data<br/>~7.5k traj, Chaos Box] --> NDM[Joint-wise neural<br/>dynamics model f_psi_i]
NDM --> Res[Residual policy<br/>a_t + a_res_t]
Gen --> Res
Res --> Pol[Action correction at deploy<br/>transfers to real]
Sim-to-real transfer for dexterous in-hand manipulation has been dominated by heavy domain randomization β train on a wide distribution and hope the real robot is in it. Works, but brittle and tuning-heavy.
Two-stage pipeline. (1) Specialist-to-generalist policy: train category-specific RL experts (PPO in IsaacGym) with privileged observations, then behavior-clone successful trajectories into a single generalist policy that uses only proprioception history, wrist orientation, and target axis. (2) Joint-wise neural dynamics model: for each joint i, learn a dynamics model q^(t+1)_i = f_Οi(h^i_t) that predicts the next joint state from only that joint's W-step state-action history (not the global hand state). This factorization contracts the high-dimensional reality gap into per-joint low-dimensional terms, making it learnable from limited real data (Claim 3.1, via data-processing inequality on KL between train/test distributions).
The model is not composed into a hybrid simulator to retrain the policy. Instead it supervises a lightweight residual policy (a_t + a^res_t) that corrects the sim-trained base policy's actions at deployment. Real data is collected autonomously via a "Chaos Box" β the hand is placed in soft balls and replays open-loop actions, yielding ~7.5k trajectories of diverse load interactions with no human resets and no object-state estimation.
A single sim-trained policy transfers to a wide range of real objects (size 2β20 cm, aspect ratios up to 5.33:1, regular/small/irregular/animal shapes) across 6 wrist orientations and multi-axis targets.
- Sim (unseen objects, Β±x axis): RotR 144.22Β±13.91 vs. AnyRotate reimpl. 91.90Β±11.60; goal-oriented success 88.27Β±3.21% vs. 64.33Β±4.70%.
- Real multi-axis (palm-down z): regular objects 23.82Β±3.86 rad rotated, 37.50Β±5.02 s time-to-fall; small objects 9.29Β±1.63 rad; irregular 8.61Β±0.76 rad.
- vs. AnyRotate (shared objects): e.g. tin cylinder 10.79Β±0.54 rad vs. 2.63Β±0.75 rad.
- Data efficiency: ~4,000 autonomous trajectories reach quality that would need ~7.5M task-aware trajectories (power-law scaling). Cross-simulator transfer (IsaacGymβGenesis, MuJoCo) outperforms UAN and ASAP.
A concrete alternative to domain randomization for sim-to-real. Rather than randomizing dynamics in sim, it measures the residual gap per joint from cheap autonomous data and corrects the policy's actions at deployment. The joint-wise factorization is the key data-efficiency lever β it provably contracts distribution shift β and the Chaos-Box autonomous collection removes the human-in-the-loop bottleneck that limits most real-world dexterous data.
- arXiv:2510.08556
- OpenReview
- Authors: Xueyi Liu (Tsinghua, Shanghai Qi Zhi), He Wang (Peking Univ., Galbot), Li Yi (Tsinghua, Shanghai Qi Zhi)
- ICLR 2026 listing
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)