-
Notifications
You must be signed in to change notification settings - Fork 0
CVPR 2026 UniDex
Venue: CVPR 2026 Category: Egocentric + Dexterous VLA Trend tag: Trend 5 (egocentric β manipulation)
flowchart LR
EGO["50k+ egocentric<br/>human-hand videos"] --> RETARGET["retarget to 8 dex hands"]
RETARGET --> FAAS["Function-Actuator-Aligned Space"]
FAAS --> VLA["3D VLA backbone"]
ROBOT["unseen dex hand"] --> VLA
VLA --> ACT["action chunk"]
Dexterous manipulation needs data per-hand because of the actuator-count mismatch (humans: ~20 DoF; robot hands: 6β24 active DoF, different topology). Existing dex datasets are tiny relative to web video. Egocentric human-hand video is abundant but cannot be directly used without bridging the actuator gap.
- Collect / curate 50 000+ retargeted trajectories (~9M paired imageβpointcloudβaction frames) across 8 dexterous hands (Inspire, Leap, Shadow, Allegro, Ability, Oymotion, Xhand, Wuji), via human-in-the-loop retargeting from egocentric human video.
- Define the Function-Actuator-Aligned Space (FAAS) β a unified action space that maps functionally similar actuators to shared coordinates regardless of the underlying actuator topology, enabling cross-hand transfer.
- Train a 3D VLA (UniDex-VLA) on the unified FAAS with multi-hand training: a Uni3D point-cloud encoder (replacing the SigLIP 2D encoder) on a PaliGemma backbone, trained with a conditional flow-matching objective.
- Ship UniDex-Cap, a portable capture setup for humanβrobot data co-training.
81 % average task progress across five real-world tool-use tasks (Make Coffee, Sweep Objects, Water Flowers, Cut Bags, Use Mouse) on two hands (Inspire, Wuji), outperforming prior VLA baselines by a large margin. Separately, a policy trained on the Inspire Hand exhibits zero-shot cross-hand transfer, reaching 60 % success on Oymotion and 40 % on Wuji without any fine-tuning, alongside spatial and object generalization.
UniDex is the dex-hand counterpart to X-VLA's soft-prompt cross-embodiment thesis: the right factorization (FAAS for dex hands; soft prompts for arms) lets one policy generalize across mechanically different end-effectors. Likely to become a baseline for subsequent dex-VLA work. Thematically related (concurrent ego-video work, different team): EgoScale (data scaling for ego video).
From Tsinghua University, Shanghai Qizhi Institute, Sun Yat-sen University, and UNC Chapel Hill.
- arXiv: 2603.22264
- Code:
unidex-ai/UniDex
- Dexterous Manipulation review Β· Cross-Embodiment review
- DexUMI Β· EgoDex Β· Human-Video Pretraining
- CVPR 2026 survey
β Back to CVPR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)