-
Notifications
You must be signed in to change notification settings - Fork 0
ICRA 2026 Dexora
Venue: ICRA 2026 Β· Authors: Zongzheng Zhang, Jingrui Pang, Zhuo Yang, Kun Li, Minwen Liao, Saining Zhang, Guoxuan Chi, et al. (Tsinghua / BAAI / PKU) Β· arXiv: 2605.18722 Category: Dexterous / bimanual VLA Trend tag: Open-source high-DoF dexterity
flowchart LR
subgraph Teleop[Hybrid teleoperation]
EXO[Exoskeleton backpack<br/>gross arm kinematics] --> PLT[Dual-arm dual-hand<br/>platform + MuJoCo twin]
AVP[Apple Vision Pro<br/>markerless finger tracking] --> PLT
end
PLT --> DATA[(Data: 100K sim<br/>+ 10K real episodes)]
DATA --> DISC[Data-quality discriminator<br/>PU objective β loss weights]
V[Multi-view RGB] --> ENC[SigLIP vision + T5 text]
L[Language] --> ENC
ENC --> BB[Decoder-only Transformer<br/>28 layers Β· 1024 hidden Β· 16 heads]
DISC -. weighted loss .-> BB
BB --> HEAD[Diffusion transformer head<br/>DDPM train Β· DPMSolver++ infer]
HEAD --> ACT[36-DoF bimanual action<br/>2Γ6 arm + 2Γ12 hand]
Most open VLAs target single-arm grippers; dexterous, bimanual, high-DoF control is gated behind closed hardware and proprietary data. Two coupled obstacles: (1) collecting high-quality demonstrations for many-DoF hands is hard, and naive teleoperation conflates gross arm motion with fine finger motion; (2) real demonstration sets are noisy, so uniform imitation wastes capacity on bad trajectories. Dexora aims to be a fully open-source stack β hardware, data, and policy β for dual-arm dual-hand dexterity.
Platform. AIRBOT 6-DoF arms paired with XHAND dexterous hands (12 fully actuated joints each), giving a 36-DoF bimanual system (2Γ6 arm + 2Γ12 hand), mirrored by an identical MuJoCo digital twin.
Hybrid teleoperation. A custom exoskeleton backpack captures gross arm kinematics while an Apple Vision Pro provides markerless finger tracking β decoupling coarse arm motion from fine hand motion so high-DoF demonstrations stay clean.
Data. Pretrain on 100K simulated bimanual-hand trajectories; post-train on 10K real teleoperated episodes (the released real-world set: 12.2K episodes, 2.92M frames, 40.5 hours; 347 objects across 17 categories).
Architecture. SigLIP encodes multi-view RGB and T5 encodes language into a decoder-only Transformer backbone (28 layers, hidden 1024, 16 heads); a diffusion-transformer head predicts action chunks (DDPM in training, DPMSolver++ for fast inference).
Discriminator-guided training. A learned data-quality discriminator scores each demonstration via a positiveβunlabeled objective; scores become per-sample weights on the diffusion loss, so the policy prioritizes high-quality trajectories and down-weights poor ones. This is a discriminator-weighted imitation variant tailored to noisy high-DoF data.
This sits in the VLM + separate diffusion action expert family (see VLA Architectures review).
On basic tasks Dexora reaches 89.6 average success, ahead of GR00T N1 (82.1), Οβ (50.4), and Diffusion Policy (34.2). On dexterous tasks it averages 66.7% vs. GR00T N1 51.7%, Οβ 26.7%, and DP 6.7%. Per-task dexterous results include Fetch Book 80%, Cut Leek 80%, Rough Dough 80%, Place Plates 70%, Use Pen 65%, Twist Cap 25%. The paper reports robust out-of-distribution and cross-embodiment generalization.
Dexora is presented as the first open-source VLA natively targeting dual-arm, dual-hand high-DoF manipulation, lowering the barrier for dexterous manipulation research. Two transferable ideas: hybrid arm/finger teleoperation for clean high-DoF data, and discriminator-weighted diffusion training that turns data-quality into a learning signal rather than relying on manual curation.
- Dexterous Manipulation review
- VLA Architectures review
- DexVLA (sibling plug-in diffusion action expert)
- ICRA 2026 Survey
β Back to ICRA-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)