Skip to content

ICRA 2026 Dexora

Heungwoo edited this page Jun 1, 2026 · 1 revision

Dexora β€” Open-Source High-DoF Bimanual Dexterous VLA

Venue: ICRA 2026 Β· Authors: Zongzheng Zhang, Jingrui Pang, Zhuo Yang, Kun Li, Minwen Liao, Saining Zhang, Guoxuan Chi, et al. (Tsinghua / BAAI / PKU) Β· arXiv: 2605.18722 Category: Dexterous / bimanual VLA Trend tag: Open-source high-DoF dexterity

Approach diagram

flowchart LR
  subgraph Teleop[Hybrid teleoperation]
    EXO[Exoskeleton backpack<br/>gross arm kinematics] --> PLT[Dual-arm dual-hand<br/>platform + MuJoCo twin]
    AVP[Apple Vision Pro<br/>markerless finger tracking] --> PLT
  end
  PLT --> DATA[(Data: 100K sim<br/>+ 10K real episodes)]
  DATA --> DISC[Data-quality discriminator<br/>PU objective β†’ loss weights]
  V[Multi-view RGB] --> ENC[SigLIP vision + T5 text]
  L[Language] --> ENC
  ENC --> BB[Decoder-only Transformer<br/>28 layers Β· 1024 hidden Β· 16 heads]
  DISC -. weighted loss .-> BB
  BB --> HEAD[Diffusion transformer head<br/>DDPM train Β· DPMSolver++ infer]
  HEAD --> ACT[36-DoF bimanual action<br/>2Γ—6 arm + 2Γ—12 hand]
Loading

Problem

Most open VLAs target single-arm grippers; dexterous, bimanual, high-DoF control is gated behind closed hardware and proprietary data. Two coupled obstacles: (1) collecting high-quality demonstrations for many-DoF hands is hard, and naive teleoperation conflates gross arm motion with fine finger motion; (2) real demonstration sets are noisy, so uniform imitation wastes capacity on bad trajectories. Dexora aims to be a fully open-source stack β€” hardware, data, and policy β€” for dual-arm dual-hand dexterity.

Method

Platform. AIRBOT 6-DoF arms paired with XHAND dexterous hands (12 fully actuated joints each), giving a 36-DoF bimanual system (2Γ—6 arm + 2Γ—12 hand), mirrored by an identical MuJoCo digital twin.

Hybrid teleoperation. A custom exoskeleton backpack captures gross arm kinematics while an Apple Vision Pro provides markerless finger tracking β€” decoupling coarse arm motion from fine hand motion so high-DoF demonstrations stay clean.

Data. Pretrain on 100K simulated bimanual-hand trajectories; post-train on 10K real teleoperated episodes (the released real-world set: 12.2K episodes, 2.92M frames, 40.5 hours; 347 objects across 17 categories).

Architecture. SigLIP encodes multi-view RGB and T5 encodes language into a decoder-only Transformer backbone (28 layers, hidden 1024, 16 heads); a diffusion-transformer head predicts action chunks (DDPM in training, DPMSolver++ for fast inference).

Discriminator-guided training. A learned data-quality discriminator scores each demonstration via a positive–unlabeled objective; scores become per-sample weights on the diffusion loss, so the policy prioritizes high-quality trajectories and down-weights poor ones. This is a discriminator-weighted imitation variant tailored to noisy high-DoF data.

This sits in the VLM + separate diffusion action expert family (see VLA Architectures review).

Results

On basic tasks Dexora reaches 89.6 average success, ahead of GR00T N1 (82.1), Ο€β‚€ (50.4), and Diffusion Policy (34.2). On dexterous tasks it averages 66.7% vs. GR00T N1 51.7%, Ο€β‚€ 26.7%, and DP 6.7%. Per-task dexterous results include Fetch Book 80%, Cut Leek 80%, Rough Dough 80%, Place Plates 70%, Use Pen 65%, Twist Cap 25%. The paper reports robust out-of-distribution and cross-embodiment generalization.

Significance

Dexora is presented as the first open-source VLA natively targeting dual-arm, dual-hand high-DoF manipulation, lowering the barrier for dexterous manipulation research. Two transferable ideas: hybrid arm/finger teleoperation for clean high-DoF data, and discriminator-weighted diffusion training that turns data-quality into a learning signal rather than relying on manual curation.

Links

Related pages

← Back to ICRA-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally