-
Notifications
You must be signed in to change notification settings - Fork 0
ICRA 2026 Topic Dexterous
Venue: IEEE ICRA 2026 Β· Vienna, Austria Β· June 1β5, 2026 Topic slice of the ICRA 2026 Survey. Papers carrying Dexterous Manipulation, In-Hand Manipulation, or Multifingered Hands as a primary keyword in the official PaperCept program.
75 papers at ICRA 2026 are tagged Dexterous Manipulation β the third-largest manipulation cluster after Reinforcement Learning (254) and Imitation Learning (197), and far larger than the explicit VLA-titled set (~46). The headline observation of this topic is that dexterity did not get absorbed by the VLA wave the way tabletop pick-and-place did. Where generalist VLAs now dominate the gripper-manipulation literature at ML venues, ICRA 2026's dexterous track remains a heterogeneous, methods-plural field where reinforcement learning in simulation, sim-to-real transfer, contact mechanics, and bespoke hand hardware are still the load-bearing techniques. This matches the analysis in the Dexterous Manipulation review: a dexterous hand is a 16β24-DoF system with sliding/rolling/re-grasping contacts, so action-chunk imitation and vanilla domain randomization do not suffice, and the field looks "more like hard RL with a really good simulator than scale-imitation-until-it-works."
What is genuinely new in the 2026 cohort: (1) cross-embodiment grasp foundation models that generate grasps for arbitrary hand morphologies at dataset scale (CEDex, MachaGrasp, T(R,O) Grasp, One-Policy-Fits-All); (2) human-video and smart-glasses data pipelines as the primary route to scaling dex demonstrations (Dexterity from Smart Lenses/AINA, DemoBot, DemoDiffusion, Deep Sensorimotor Control); (3) the first native dexterous/bimanual VLAs entering the venue, anchored by Dexora; and (4) a continued, vigorous hardware track β anthropomorphic, tendon-driven, underactuated, suction-augmented, and reconfigurable hands. Tactile-driven dexterity and contact-rich model-based control round out the picture as the venue's systems-and-mechanics signature.
Reinforcement learning in simulation remains the backbone for high-DoF in-hand skills, and the 2026 papers concentrate on closing the transfer gap rather than re-deriving the controller. DexCtrl: Towards Sim-to-Real Dexterity with Adaptive Controller Learning (arXiv 2505.00991) makes the sharpest argument: the dominant sim-to-real obstacle is low-level controller mismatch β identical trajectories produce different contact forces under different control parameters β so DexCtrl jointly learns actions and controller parameters and adapts gains online, sidestepping excessive randomization. Best of Sim and Real (WeI2I.101) decouples the stack entirely, learning control in simulation with privileged state and perception in the real world. DexSinGrasp (arXiv 2504.04516) trains a unified RL policy for object singulation-then-grasping in dense clutter with a clutter-arrangement curriculum, then distills to a vision policy for deployment. Deformable Cluster Manipulation via Whole-Arm Policy Learning (WeI1I.411) and Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL (ThI2I.377, applied to non-prehensile manipulation) extend RL into deformables and exploration-hard regimes. The throughline matches the review's thesis: dexterous RL in 2026 is about learned corrections to the transfer problem, not brute-force DR.
The most crowded analytical cluster is scalable, embodiment-agnostic dexterous grasp generation. CEDex (arXiv 2509.24661, code released) builds what the authors call the largest cross-embodiment grasp dataset β 500K objects, four gripper types, 20M grasps β by aligning robot kinematics to human-like contact representations via topological merging plus SDF-based optimization. MachaGrasp (arXiv 2510.06068) takes an eigengrasp route: a morphology embedding conditions an amplitude predictor that regresses low-dimensional articulation coefficients, hitting 91.9% sim success on unseen objects across three hands with sub-0.4 s inference. T(R,O) Grasp (arXiv 2510.12724) frames grasping as graph diffusion over robotβobject spatial transformations, reporting 94.83% average success at 41 grasps/s. UltraDexGrasp (arXiv 2603.05312, InternRobotics) curates UltraDexGrasp-20M (20M frames, 1000 objects) for bimanual dexterous grasping and reaches 81.2% real-world zero-shot. Complementary entries: One-Policy-Fits-All (TuI2I.169, geometry-aware action latents for cross-embodiment), Adversarial Game-Theoretic Algorithm for Dexterous Grasp Synthesis (ThI2I.235), CoorGrasp (TuBT3.3, coordinated contact control under uncertainty), and the Multi-Level Similarity single-view grasping pipeline (WeI1I.36). This is the part of the field that is converging on a foundation-model paradigm β but for grasp synthesis, not closed-loop policy.
A distinct sub-theme moves past stability toward functional grasps β grasping a spray bottle so you can actuate it, a cup by its handle, a tool for use. CoDex (ThI1I.142) studies Compositional Dexterous Functional Object Manipulation, discovering grasp-move-actuate sequences from VLM-generated semantic goals plus RL, without demonstrations. Generate, Transfer, Adapt (arXiv 2601.05243, "CorDex") learns functional grasps of novel category instances from a single human demonstration via a correspondence-based data engine. Multi-Keypoint Affordance Representation (arXiv 2502.20018) introduces Contact-guided Multi-Keypoint Affordance with weak supervision from human grasp images. UniFucGrasp (arXiv 2508.03339) contributes the first multi-hand functional grasp dataset and annotation strategy, deliberately avoiding bulky Shadow-Hand-only pipelines. Language-Guided Dexterous Functional Grasping (TuAT3.9, LLM-generated functionality and synergy for humanoids) and OmniDexGrasp (arXiv 2510.23119) β which uses foundation models to generate human grasp images and transfers them to robot actions with force-aware adaptation β close the loop between semantics and execution.
Because dex demonstrations are expensive, a large 2026 cluster scales data from humans in the wild. Dexterity from Smart Lenses (arXiv 2511.16661, framework "AINA") collects multi-fingered policy data from anyone, anywhere using Aria Gen 2 glasses, lifting human video to approximate 4D (hand keypoints + stereo depth + object pointclouds) and learning 3D keypoint policies to minimize the humanβrobot gap. DemoBot (arXiv 2601.01651, ByteDance Seed / Edinburgh) learns bimanual dexterous skills from a single unannotated RGB-D human video, using extracted hand/object trajectories as motion priors for a temporal-segment RL pipeline. DemoDiffusion (arXiv 2506.20668) retargets one human demo into a coarse robot trajectory and refines it with a pre-trained generalist diffusion policy β no online RL, no paired data. Deep Sensorimotor Control by Imitating Predictive Models of Human Motion (arXiv 2508.18691, Berkeley/Penn) trains policies to track a human-motion predictor applied zero-shot to robot keypoints, bypassing kinematic retargeting and adversarial losses. Supporting infrastructure: Robust Hand Tracking from Visual-Inertial Fusion (ThI2I.231) and Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt (TuI2I.173).
Touch is the venue's signature modality, and several papers ground dexterity in tactile geometry. Spatially-Anchored Tactile Awareness (SaTA) (arXiv 2510.14647) anchors tactile features to the hand's kinematic frame via forward kinematics, enabling sub-millimeter tasks (USB-C mating, light-bulb threading) with up to +30% success over visuo-tactile baselines. FBI / Flow before Imitation (arXiv 2508.14441) establishes a causal link between tactile signals and object motion through a dynamics-aware latent model, fusing flow-derived tactile features with vision in a one-step diffusion policy for in-hand tasks. Touch-Based Object Localisation with Spatially-Aware Belief Entropy (WeI1I.109), Hydrosoft (TuI1I.169, non-holonomic hydroelastic models for compliant tactile manipulation), and Bi-Hap (WeI1I.165, bidirectional haptic feedback for in-hand telemanipulation) extend the theme into localization, compliant contact modeling, and teleoperation feedback.
ICRA's mechanical-design heritage shows in a deep hardware track. New hands include the bioinspired tendon-driven DexHand 021 with proprioceptive compliance (ThI2I.352), MultiHand (multi-modal grasping, ThI2I.125), the underactuated anthropomorphic SoftHand Model-W with integrated wrist/carpal tunnel (TuI1I.206), an Adaptive Modular Anthropomorphic Dexterous Hand (TuI1I.67), The Folding Hand with a reconfigurable humanoid palm (TuI2I.34), the Ultra-Low-Impedance 9-DoF direct-drive gripper (WeI2LB.10), and a Wire-Driven Hand with Mode-Switchable Planetary Transmission for dynamic manipulation (TuI1LB.4). The standout conceptual hardware contribution is the Suction Leap-Hand (arXiv 2509.20646), which augments each fingertip with a suction cup to replace force-closure with single-point adhesion β simplifying teleoperation and unlocking skills humans cannot do (one-handed paper cutting, in-hand writing). Teleoperation/retargeting work: TransDexNet (ThI2I.296, end-to-end RGBβdex retargeting transformer), Switchable Neural Teleoperation (ThI1I.56), and MiniBEE (TuI1I.292, compact bimanual form factor).
The bridge to the VLA literature is anchored by Dexora (arXiv 2605.18722) β the first fully open-source 36-DoF dual-arm/dual-hand VLA, combining hybrid exoskeleton+Vision-Pro teleoperation, 100K sim + 10K real episodes, and discriminator-weighted diffusion training; it reaches 66.7% dexterous-task success vs GR00T N1's 51.7%. DextrAH-RGB (arXiv 2412.01791, NVIDIA) distills a privileged fabric-guided RL policy into an end-to-end RGB visuomotor grasping policy with photorealistic sim, claiming the first robust sim2real of an end-to-end RGB policy for contact-rich dex grasping. DQ-RISE / Learning Dexterous Manipulation with Quantized Hand State (arXiv 2509.17450) quantizes hand state so high-dimensional hand actions don't dominate the coupled armβhand action space, then diffuses arm actions jointly. SEM (ThI2I.143, 3D spatial understanding), VFP (arXiv, variational flow-matching for multi-modal manipulation, TuI1I.183), and Multi-Modal Affordance Planner for long-horizon bimanual tasks (TuI1I.79) extend generalist policy learning toward dex/bimanual settings. Consistent with the review's thesis, even the generalist entries keep RL or privileged-state distillation as the internal engine.
A rigorous model-based / planning cluster targets the combinatorics of contact. Irrotational Contact Fields (arXiv 2312.03908, Toyota Research Institute) generates convex approximations of experimentally-validated contact models (Hunt & Crossley + Coulomb friction), implemented differentiably in Drake. Approximating Global Contact-Implicit MPC via Sampling and Local Complementarity (TuI1I.375) and Push Anything (WeAT1.5, contact-implicit MPC for single/multi-object pushing) attack real-time global contact reasoning; Approximated Collision Detection for Contact-Rich Dexterous Manipulation (TuI2I.95, GJK-EPA-based NNLS), Spectral Decomposition of Inverse Dynamics (TuI2I.337), and Diffusing Trajectory Optimization Problems for Recovery (TuI2I.218) supply the supporting machinery. A non-prehensile/dynamic sub-cluster β High-Speed Scooping (ThI2I.26), Dynamic Scoop-and-Flick (TuI1I.124), Robotic Throwing / Sliding Pivot Model (WeI1I.386), Dexterous Planar Pushing under Uncertainty (TuI1I.198), and friction-modulation fingertips with passive rollers (WeAT1.2) β rounds out the contact-mechanics frontier.
These papers I confirmed with arXiv ID plus at least one reported number.
-
Dexora β Open-source 36-DoF bimanual dexterous VLA (arXiv 2605.18722). AIRBOT arms + XHAND hands (2Γ6 arm + 2Γ12 hand = 36 DoF); 100K sim + 10K real episodes; discriminator-weighted diffusion. 89.6 avg success on basic tasks and 66.7% on dexterous tasks vs GR00T N1's 51.7%. The reference open dex-VLA stack β see dedicated page.
-
CEDex β Cross-embodiment dexterous grasp generation at scale (arXiv 2509.24661, code). Aligns robot kinematics to human-like contact representations (topological merging + SDF optimization). Builds 500K objects Β· 4 grippers Β· 20M grasps, presented as the largest cross-embodiment grasp dataset. The clearest evidence that grasp synthesis (unlike closed-loop policy) is converging on a foundation-model paradigm.
-
DexCtrl β Sim-to-real via adaptive controller learning (arXiv 2505.00991). Reframes the sim-to-real gap as controller-dynamics mismatch and jointly learns actions + controller parameters with online adaptation. 85% real-world success, β40% task time vs baselines. The cleanest articulation of the review's "learned corrections beat brute DR" thesis.
-
Dexterity from Smart Lenses (AINA) (arXiv 2511.16661). Multi-fingered policies from in-the-wild human video captured with Aria Gen 2 glasses; lifts video to approximate 4D (hand keypoints + stereo depth + object pointclouds) and learns 3D-keypoint policies to minimize the humanβrobot domain gap. The most ambitious "anyone, anywhere" dex data-collection proposal in the cohort.
-
Suction Leap-Hand (arXiv 2509.20646). Augments each LEAP fingertip with a suction cup, replacing force-closure with single-point adhesion. Stabilizes teleoperation and the sim-to-real gap, and unlocks non-human skills (one-handed paper cutting, in-hand writing). A rare hardware paper that changes the learning problem, not just the mechanism.
-
DextrAH-RGB β End-to-end RGB dexterous grasping (arXiv 2412.01791, NVIDIA). Distills a privileged fabric-guided RL policy into an RGB visuomotor policy via photorealistic tiled rendering; claimed first robust sim2real of an end-to-end RGB policy for contact-rich dex grasping, competitive with depth-based policies and generalizing to unseen geometry/texture/lighting.
-
SaTA β Spatially-anchored tactile awareness (arXiv 2510.14647). Anchors tactile features to the hand frame via forward kinematics for sub-millimeter tasks (USB-C mating, light-bulb threading). Reports up to +30% success and β27% completion time over visuo-tactile baselines β the strongest tactile-dexterity result in the set.
| Code | Title | arXiv |
|---|---|---|
| ThBT2.4 | Deep Sensorimotor Control by Imitating Predictive Models of Human Motion | 2508.18691 |
| ThI1I.107 | DemoBot: Efficient Learning of Bimanual Manipulation with Dexterous Hands from Third-Person Human Videos | 2601.01651 |
| ThI1I.142 | CoDex: Learning Compositional Dexterous Functional Manipulation without Demonstrations | β |
| ThI1I.244 | In-The-Wild Compliant Manipulation with UMI-FT | 2601.09988 |
| ThI1I.320 | MachaGrasp: Morphology-Aware Cross-Embodiment Dexterous Hand Articulation Generation for Grasping | 2510.06068 |
| ThI1I.321 | OmniDexGrasp: Generalizable Dexterous Grasping via Foundation Model and Force Feedback | 2510.23119 |
| ThI1I.346 | Multi-Keypoint Affordance Representation for Functional Dexterous Grasping | 2502.20018 |
| ThI1I.381 | Robot Deformable Object Manipulation via NMPC-Generated Demonstrations in Deep RL (I) | 2502.11375 |
| ThI1I.386 | The Developments and Challenges towards Dexterous and Embodied Robotic Manipulation: A Survey | 2507.11840 |
| ThI1I.56 | Switchable Neural Teleoperation | β |
| ThI1I.85 | Monorail-Like Gripper System with Dynamic and Modular Reconfiguration for Diverse Finger Layouts | β |
| ThI1I.96 | DemoDiffusion: One-Shot Human Imitation Using Pre-Trained Diffusion Policy | 2506.20668 |
| ThI2I.125 | MultiHand: Design and Verification of a Dexterous Hand with Multi-Modal Grasping Capabilities | β |
| ThI2I.143 | SEM: Enhancing Spatial Understanding for Robust Robot Manipulation | 2505.16196 |
| ThI2I.231 | Robust Hand Tracking from Visual-Inertial Fusion | β |
| ThI2I.235 | Adversarial Game-Theoretic Algorithm for Dexterous Grasp Synthesis | 2511.05809 |
| ThI2I.26 | High-Speed Scooping through Dynamic Manipulation: Model and Practice | β |
| ThI2I.292 | Generate, Transfer, Adapt: Learning Functional Dexterous Grasping from a Single Human Demonstration | 2601.05243 |
| ThI2I.296 | TransDexNet: End-to-End Motion Retargeting Network with Transformer for Dexterous Hand Teleoperation from RGB Images | β |
| ThI2I.352 | Development of the Bioinspired Tendon-Driven DexHand 021 with Proprioceptive Compliance Control | 2511.03481 |
| ThI2I.377 | Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation | 2411.14913 |
| ThI2I.383 | DexSinGrasp: Learning a Unified Policy for Dexterous Object Singulation and Grasping in Densely Cluttered Environments | 2504.04516 |
| ThI2I.83 | FAR-Dex: Few-Shot Data Augmentation and Adaptive Residual Policy Refinement for Dexterous Manipulation | 2603.10451 |
| TuAT3.1 | Spatially-Anchored Tactile Awareness for Robust Dexterous Manipulation (SaTA) | 2510.14647 |
| TuAT3.7 | Irrotational Contact Fields | 2312.03908 |
| TuAT3.9 | Language-Guided Dexterous Functional Grasping by LLM Generated Grasp Functionality and Synergy for Humanoid Manipulation (I) | β |
| TuAT4.6 | Soft Omni-Functional Robotic Gripper with a Force-Enhanced Pleated Mechanism for High Force and Multi-DoF Manipulation | β |
| TuBT3.3 | CoorGrasp: Coordinated Contact Control for Adaptive Dexterous Grasping under Uncertainty | β |
| TuI1I.124 | Dynamic Scoop-and-Flick Manipulation for Rapid Non-Prehensile High-Arc Object Transfer | β |
| TuI1I.169 | Hydrosoft: Non-Holonomic Hydroelastic Models for Compliant Tactile Manipulation | 2509.13126 |
| TuI1I.171 | CEDex: Cross-Embodiment Dexterous Grasp Generation at Scale from Human-Like Contact Representations | 2509.24661 |
| TuI1I.183 | VFP: Variational Flow-Matching Policy for Multi-Modal Robot Manipulation | 2508.01622 |
| TuI1I.198 | Dexterous Planar Pushing under Uncertain Object Properties: A Contact-Aware Goal-Oriented Approach | β |
| TuI1I.199 | DexKnot: Generalizable Visuomotor Policy Learning for Dexterous Bag-Knotting Manipulation | 2603.07136 |
| TuI1I.206 | SoftHand Model-W: A 3D-Printed, Anthropomorphic, Underactuated Robot Hand with Integrated Wrist and Carpal Tunnel | 2604.00738 |
| TuI1I.214 | Two Degree-of-Freedom Vibratory Transport in a Grasp | β |
| TuI1I.249 | Suction Leap-Hand: Suction Cups on a Multi-Fingered Hand Enable Embodied Dexterity and In-Hand Teleoperation | 2509.20646 |
| TuI1I.292 | MiniBEE: A New Form Factor for Compact Bimanual Dexterity | 2510.01603 |
| TuI1I.375 | Approximating Global Contact-Implicit MPC via Sampling and Local Complementarity | 2505.13350 |
| TuI1I.52 | A Gripper with Extreme Stiffness Anisotropy for High-Speed Handling of Fragile Foods | β |
| TuI1I.67 | Design of an Adaptive Modular Anthropomorphic Dexterous Hand for Human-Like Manipulation | 2511.22100 |
| TuI1I.79 | Multi-Modal Affordance Planner with Temporal-Context Action Policy for Long-Horizon Bimanual Robot Manipulation | β |
| TuI1LB.4 | A Wire-Driven Robotic Hand with Mode-Switchable Planetary Transmission for Dynamic Manipulation | β |
| TuI2I.11 | Nonlinear Model Predictive Control for Robotic Pushing of Planar Objects with Generic Shape | β |
| TuI2I.111 | Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-The-Wild Human Demonstrations (AINA) | 2511.16661 |
| TuI2I.141 | DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands | 2412.01791 |
| TuI2I.169 | One-Policy-Fits-All: Geometry-Aware Action Latents for Cross-Embodiment Manipulation | 2603.14522 |
| TuI2I.173 | Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt | 2505.20795 |
| TuI2I.196 | UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data | 2603.05312 |
| TuI2I.218 | Diffusing Trajectory Optimization Problems for Recovery During Multi-Finger Manipulation | 2510.07030 |
| TuI2I.309 | Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation | 2510.08807 |
| TuI2I.337 | Spectral Decomposition of Inverse Dynamics for Fast Exploration in Model-Based Manipulation | 2603.27796 |
| TuI2I.34 | The Folding Hand: Anthropomorphic Robotic Hands with a Compact Reconfigurable Humanoid Palm Design | β |
| TuI2I.67 | Learning Geometry-Aware Nonprehensile Pushing and Pulling with Dexterous Hands | 2509.18455 |
| TuI2I.95 | Approximated Collision Detection for Contact-Rich Dexterous Manipulation with Nonnegative Least Squares | β |
| WeAT1.2 | Robotic Dexterous Manipulation via Anisotropic Friction Modulation Using Passive Rollers | 2603.27452 |
| WeAT1.5 | Push Anything: Single and Multi-Object Pushing from First Sight with Contact-Implicit MPC | 2510.19974 |
| WeI1I.109 | Touch-Based Object Localisation with Spatially-Aware Belief Entropy Estimation | β |
| WeI1I.165 | Bi-Hap: A Bi-Directional Learning-Based Control and Momentum-Based Haptic Feedback System for Dexterous In-Hand Telemanipulation | 2409.20527 |
| WeI1I.183 | HybNetic: A Mobile Hybrid Magnetic Actuation System | β |
| WeI1I.213 | DexCtrl: Sim-To-Real Dexterity with Adaptive Controller Learning | 2505.00991 |
| WeI1I.345 | Frictional and Prismatic Pin-Array Gripper for Universal Gripping and Stable Tool Manipulation | β |
| WeI1I.36 | A Multi-Level Similarity Approach for Single-View Object Grasping: Matching, Planning, and Fine-Tuning | 2507.11938 |
| WeI1I.376 | UniFucGrasp: Human-Hand-Inspired Unified Functional Grasp Annotation Strategy and Dataset for Diverse Dexterous Hands | 2508.03339 |
| WeI1I.386 | On Transient Release Dynamics in Robot Throwing: A Sliding Pivot Model | β |
| WeI1I.411 | Deformable Cluster Manipulation via Whole-Arm Policy Learning | 2507.17085 |
| WeI2I.101 | Best of Sim and Real: Decoupled Visuomotor Manipulation via Learning Control in Simulation and Perception in Real | 2509.25747 |
| WeI2I.13 | Enhancing Safety and Manipulability of Redundant Manipulators: Accelerated Motion Generation in Dynamic Environments | β |
| WeI2I.176 | T(R,O) Grasp: Efficient Graph Diffusion of Robot-Object Spatial Transformation for Cross-Embodiment Dexterous Grasping | 2510.12724 |
| WeI2I.204 | Learning Dexterous Manipulation with Quantized Hand State (DQ-RISE) | 2509.17450 |
| WeI2I.292 | Towards Automated Chicken Deboning via Learning-Based Dynamically-Adaptive 6-DoF Multi-Material Cutting | 2510.15376 |
| WeI2I.338 | Manipulating Elasto-Plastic Objects with 3D Occupancy and Learning-Based Predictive Control | 2505.16249 |
| WeI2I.428 | Estimating Deformable-Rigid Contact Interactions for a Deformable Tool via Learning and Model-Based Optimization | 2505.10884 |
| WeI2I.55 | Flow before Imitation: Learning Dexterous In-Hand Manipulation with Dynamic Visuotactile Shortcut Policy (FBI) | 2508.14441 |
| WeI2LB.10 | Ultra-Low-Impedance Robotic Gripper for High-Bandwidth and Transparent Physical Interaction | β |
β Back to ICRA-2026-VLA-Manipulation-Survey
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)