Skip to content

ICRA 2026 Topic Dexterous

Heungwoo edited this page Jun 1, 2026 · 2 revisions

ICRA 2026 β€” Dexterous & In-Hand Manipulation (Topic Analysis)

Venue: IEEE ICRA 2026 Β· Vienna, Austria Β· June 1–5, 2026 Topic slice of the ICRA 2026 Survey. Papers carrying Dexterous Manipulation, In-Hand Manipulation, or Multifingered Hands as a primary keyword in the official PaperCept program.

Intro

75 papers at ICRA 2026 are tagged Dexterous Manipulation β€” the third-largest manipulation cluster after Reinforcement Learning (254) and Imitation Learning (197), and far larger than the explicit VLA-titled set (~46). The headline observation of this topic is that dexterity did not get absorbed by the VLA wave the way tabletop pick-and-place did. Where generalist VLAs now dominate the gripper-manipulation literature at ML venues, ICRA 2026's dexterous track remains a heterogeneous, methods-plural field where reinforcement learning in simulation, sim-to-real transfer, contact mechanics, and bespoke hand hardware are still the load-bearing techniques. This matches the analysis in the Dexterous Manipulation review: a dexterous hand is a 16–24-DoF system with sliding/rolling/re-grasping contacts, so action-chunk imitation and vanilla domain randomization do not suffice, and the field looks "more like hard RL with a really good simulator than scale-imitation-until-it-works."

What is genuinely new in the 2026 cohort: (1) cross-embodiment grasp foundation models that generate grasps for arbitrary hand morphologies at dataset scale (CEDex, MachaGrasp, T(R,O) Grasp, One-Policy-Fits-All); (2) human-video and smart-glasses data pipelines as the primary route to scaling dex demonstrations (Dexterity from Smart Lenses/AINA, DemoBot, DemoDiffusion, Deep Sensorimotor Control); (3) the first native dexterous/bimanual VLAs entering the venue, anchored by Dexora; and (4) a continued, vigorous hardware track β€” anthropomorphic, tendon-driven, underactuated, suction-augmented, and reconfigurable hands. Tactile-driven dexterity and contact-rich model-based control round out the picture as the venue's systems-and-mechanics signature.

Sub-trends

1. RL for in-hand / contact-rich dexterity and the sim-to-real gap

Reinforcement learning in simulation remains the backbone for high-DoF in-hand skills, and the 2026 papers concentrate on closing the transfer gap rather than re-deriving the controller. DexCtrl: Towards Sim-to-Real Dexterity with Adaptive Controller Learning (arXiv 2505.00991) makes the sharpest argument: the dominant sim-to-real obstacle is low-level controller mismatch β€” identical trajectories produce different contact forces under different control parameters β€” so DexCtrl jointly learns actions and controller parameters and adapts gains online, sidestepping excessive randomization. Best of Sim and Real (WeI2I.101) decouples the stack entirely, learning control in simulation with privileged state and perception in the real world. DexSinGrasp (arXiv 2504.04516) trains a unified RL policy for object singulation-then-grasping in dense clutter with a clutter-arrangement curriculum, then distills to a vision policy for deployment. Deformable Cluster Manipulation via Whole-Arm Policy Learning (WeI1I.411) and Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL (ThI2I.377, applied to non-prehensile manipulation) extend RL into deformables and exploration-hard regimes. The throughline matches the review's thesis: dexterous RL in 2026 is about learned corrections to the transfer problem, not brute-force DR.

2. Dexterous grasping foundation / cross-embodiment generation models

The most crowded analytical cluster is scalable, embodiment-agnostic dexterous grasp generation. CEDex (arXiv 2509.24661, code released) builds what the authors call the largest cross-embodiment grasp dataset β€” 500K objects, four gripper types, 20M grasps β€” by aligning robot kinematics to human-like contact representations via topological merging plus SDF-based optimization. MachaGrasp (arXiv 2510.06068) takes an eigengrasp route: a morphology embedding conditions an amplitude predictor that regresses low-dimensional articulation coefficients, hitting 91.9% sim success on unseen objects across three hands with sub-0.4 s inference. T(R,O) Grasp (arXiv 2510.12724) frames grasping as graph diffusion over robot–object spatial transformations, reporting 94.83% average success at 41 grasps/s. UltraDexGrasp (arXiv 2603.05312, InternRobotics) curates UltraDexGrasp-20M (20M frames, 1000 objects) for bimanual dexterous grasping and reaches 81.2% real-world zero-shot. Complementary entries: One-Policy-Fits-All (TuI2I.169, geometry-aware action latents for cross-embodiment), Adversarial Game-Theoretic Algorithm for Dexterous Grasp Synthesis (ThI2I.235), CoorGrasp (TuBT3.3, coordinated contact control under uncertainty), and the Multi-Level Similarity single-view grasping pipeline (WeI1I.36). This is the part of the field that is converging on a foundation-model paradigm β€” but for grasp synthesis, not closed-loop policy.

3. Functional / tool-use dexterity (affordance + semantics)

A distinct sub-theme moves past stability toward functional grasps β€” grasping a spray bottle so you can actuate it, a cup by its handle, a tool for use. CoDex (ThI1I.142) studies Compositional Dexterous Functional Object Manipulation, discovering grasp-move-actuate sequences from VLM-generated semantic goals plus RL, without demonstrations. Generate, Transfer, Adapt (arXiv 2601.05243, "CorDex") learns functional grasps of novel category instances from a single human demonstration via a correspondence-based data engine. Multi-Keypoint Affordance Representation (arXiv 2502.20018) introduces Contact-guided Multi-Keypoint Affordance with weak supervision from human grasp images. UniFucGrasp (arXiv 2508.03339) contributes the first multi-hand functional grasp dataset and annotation strategy, deliberately avoiding bulky Shadow-Hand-only pipelines. Language-Guided Dexterous Functional Grasping (TuAT3.9, LLM-generated functionality and synergy for humanoids) and OmniDexGrasp (arXiv 2510.23119) β€” which uses foundation models to generate human grasp images and transfers them to robot actions with force-aware adaptation β€” close the loop between semantics and execution.

4. Human-video, glove & smart-glasses data for dexterity

Because dex demonstrations are expensive, a large 2026 cluster scales data from humans in the wild. Dexterity from Smart Lenses (arXiv 2511.16661, framework "AINA") collects multi-fingered policy data from anyone, anywhere using Aria Gen 2 glasses, lifting human video to approximate 4D (hand keypoints + stereo depth + object pointclouds) and learning 3D keypoint policies to minimize the human–robot gap. DemoBot (arXiv 2601.01651, ByteDance Seed / Edinburgh) learns bimanual dexterous skills from a single unannotated RGB-D human video, using extracted hand/object trajectories as motion priors for a temporal-segment RL pipeline. DemoDiffusion (arXiv 2506.20668) retargets one human demo into a coarse robot trajectory and refines it with a pre-trained generalist diffusion policy β€” no online RL, no paired data. Deep Sensorimotor Control by Imitating Predictive Models of Human Motion (arXiv 2508.18691, Berkeley/Penn) trains policies to track a human-motion predictor applied zero-shot to robot keypoints, bypassing kinematic retargeting and adversarial losses. Supporting infrastructure: Robust Hand Tracking from Visual-Inertial Fusion (ThI2I.231) and Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt (TuI2I.173).

5. Tactile- and contact-driven dexterity

Touch is the venue's signature modality, and several papers ground dexterity in tactile geometry. Spatially-Anchored Tactile Awareness (SaTA) (arXiv 2510.14647) anchors tactile features to the hand's kinematic frame via forward kinematics, enabling sub-millimeter tasks (USB-C mating, light-bulb threading) with up to +30% success over visuo-tactile baselines. FBI / Flow before Imitation (arXiv 2508.14441) establishes a causal link between tactile signals and object motion through a dynamics-aware latent model, fusing flow-derived tactile features with vision in a one-step diffusion policy for in-hand tasks. Touch-Based Object Localisation with Spatially-Aware Belief Entropy (WeI1I.109), Hydrosoft (TuI1I.169, non-holonomic hydroelastic models for compliant tactile manipulation), and Bi-Hap (WeI1I.165, bidirectional haptic feedback for in-hand telemanipulation) extend the theme into localization, compliant contact modeling, and teleoperation feedback.

6. Multifingered hand hardware, teleoperation & retargeting

ICRA's mechanical-design heritage shows in a deep hardware track. New hands include the bioinspired tendon-driven DexHand 021 with proprioceptive compliance (ThI2I.352), MultiHand (multi-modal grasping, ThI2I.125), the underactuated anthropomorphic SoftHand Model-W with integrated wrist/carpal tunnel (TuI1I.206), an Adaptive Modular Anthropomorphic Dexterous Hand (TuI1I.67), The Folding Hand with a reconfigurable humanoid palm (TuI2I.34), the Ultra-Low-Impedance 9-DoF direct-drive gripper (WeI2LB.10), and a Wire-Driven Hand with Mode-Switchable Planetary Transmission for dynamic manipulation (TuI1LB.4). The standout conceptual hardware contribution is the Suction Leap-Hand (arXiv 2509.20646), which augments each fingertip with a suction cup to replace force-closure with single-point adhesion — simplifying teleoperation and unlocking skills humans cannot do (one-handed paper cutting, in-hand writing). Teleoperation/retargeting work: TransDexNet (ThI2I.296, end-to-end RGB→dex retargeting transformer), Switchable Neural Teleoperation (ThI1I.56), and MiniBEE (TuI1I.292, compact bimanual form factor).

7. Dexterous VLA & generalist policies

The bridge to the VLA literature is anchored by Dexora (arXiv 2605.18722) β€” the first fully open-source 36-DoF dual-arm/dual-hand VLA, combining hybrid exoskeleton+Vision-Pro teleoperation, 100K sim + 10K real episodes, and discriminator-weighted diffusion training; it reaches 66.7% dexterous-task success vs GR00T N1's 51.7%. DextrAH-RGB (arXiv 2412.01791, NVIDIA) distills a privileged fabric-guided RL policy into an end-to-end RGB visuomotor grasping policy with photorealistic sim, claiming the first robust sim2real of an end-to-end RGB policy for contact-rich dex grasping. DQ-RISE / Learning Dexterous Manipulation with Quantized Hand State (arXiv 2509.17450) quantizes hand state so high-dimensional hand actions don't dominate the coupled arm–hand action space, then diffuses arm actions jointly. SEM (ThI2I.143, 3D spatial understanding), VFP (arXiv, variational flow-matching for multi-modal manipulation, TuI1I.183), and Multi-Modal Affordance Planner for long-horizon bimanual tasks (TuI1I.79) extend generalist policy learning toward dex/bimanual settings. Consistent with the review's thesis, even the generalist entries keep RL or privileged-state distillation as the internal engine.

8. Contact-implicit & model-based control for contact-rich manipulation

A rigorous model-based / planning cluster targets the combinatorics of contact. Irrotational Contact Fields (arXiv 2312.03908, Toyota Research Institute) generates convex approximations of experimentally-validated contact models (Hunt & Crossley + Coulomb friction), implemented differentiably in Drake. Approximating Global Contact-Implicit MPC via Sampling and Local Complementarity (TuI1I.375) and Push Anything (WeAT1.5, contact-implicit MPC for single/multi-object pushing) attack real-time global contact reasoning; Approximated Collision Detection for Contact-Rich Dexterous Manipulation (TuI2I.95, GJK-EPA-based NNLS), Spectral Decomposition of Inverse Dynamics (TuI2I.337), and Diffusing Trajectory Optimization Problems for Recovery (TuI2I.218) supply the supporting machinery. A non-prehensile/dynamic sub-cluster β€” High-Speed Scooping (ThI2I.26), Dynamic Scoop-and-Flick (TuI1I.124), Robotic Throwing / Sliding Pivot Model (WeI1I.386), Dexterous Planar Pushing under Uncertainty (TuI1I.198), and friction-modulation fingertips with passive rollers (WeAT1.2) β€” rounds out the contact-mechanics frontier.

Standout deep-dives

These papers I confirmed with arXiv ID plus at least one reported number.

  • Dexora β€” Open-source 36-DoF bimanual dexterous VLA (arXiv 2605.18722). AIRBOT arms + XHAND hands (2Γ—6 arm + 2Γ—12 hand = 36 DoF); 100K sim + 10K real episodes; discriminator-weighted diffusion. 89.6 avg success on basic tasks and 66.7% on dexterous tasks vs GR00T N1's 51.7%. The reference open dex-VLA stack β€” see dedicated page.

  • CEDex β€” Cross-embodiment dexterous grasp generation at scale (arXiv 2509.24661, code). Aligns robot kinematics to human-like contact representations (topological merging + SDF optimization). Builds 500K objects Β· 4 grippers Β· 20M grasps, presented as the largest cross-embodiment grasp dataset. The clearest evidence that grasp synthesis (unlike closed-loop policy) is converging on a foundation-model paradigm.

  • DexCtrl β€” Sim-to-real via adaptive controller learning (arXiv 2505.00991). Reframes the sim-to-real gap as controller-dynamics mismatch and jointly learns actions + controller parameters with online adaptation. 85% real-world success, βˆ’40% task time vs baselines. The cleanest articulation of the review's "learned corrections beat brute DR" thesis.

  • Dexterity from Smart Lenses (AINA) (arXiv 2511.16661). Multi-fingered policies from in-the-wild human video captured with Aria Gen 2 glasses; lifts video to approximate 4D (hand keypoints + stereo depth + object pointclouds) and learns 3D-keypoint policies to minimize the human–robot domain gap. The most ambitious "anyone, anywhere" dex data-collection proposal in the cohort.

  • Suction Leap-Hand (arXiv 2509.20646). Augments each LEAP fingertip with a suction cup, replacing force-closure with single-point adhesion. Stabilizes teleoperation and the sim-to-real gap, and unlocks non-human skills (one-handed paper cutting, in-hand writing). A rare hardware paper that changes the learning problem, not just the mechanism.

  • DextrAH-RGB β€” End-to-end RGB dexterous grasping (arXiv 2412.01791, NVIDIA). Distills a privileged fabric-guided RL policy into an RGB visuomotor policy via photorealistic tiled rendering; claimed first robust sim2real of an end-to-end RGB policy for contact-rich dex grasping, competitive with depth-based policies and generalizing to unseen geometry/texture/lighting.

  • SaTA β€” Spatially-anchored tactile awareness (arXiv 2510.14647). Anchors tactile features to the hand frame via forward kinematics for sub-millimeter tasks (USB-C mating, light-bulb threading). Reports up to +30% success and βˆ’27% completion time over visuo-tactile baselines β€” the strongest tactile-dexterity result in the set.

Complete paper list (75)

Code Title arXiv
ThBT2.4 Deep Sensorimotor Control by Imitating Predictive Models of Human Motion 2508.18691
ThI1I.107 DemoBot: Efficient Learning of Bimanual Manipulation with Dexterous Hands from Third-Person Human Videos 2601.01651
ThI1I.142 CoDex: Learning Compositional Dexterous Functional Manipulation without Demonstrations β€”
ThI1I.244 In-The-Wild Compliant Manipulation with UMI-FT 2601.09988
ThI1I.320 MachaGrasp: Morphology-Aware Cross-Embodiment Dexterous Hand Articulation Generation for Grasping 2510.06068
ThI1I.321 OmniDexGrasp: Generalizable Dexterous Grasping via Foundation Model and Force Feedback 2510.23119
ThI1I.346 Multi-Keypoint Affordance Representation for Functional Dexterous Grasping 2502.20018
ThI1I.381 Robot Deformable Object Manipulation via NMPC-Generated Demonstrations in Deep RL (I) 2502.11375
ThI1I.386 The Developments and Challenges towards Dexterous and Embodied Robotic Manipulation: A Survey 2507.11840
ThI1I.56 Switchable Neural Teleoperation β€”
ThI1I.85 Monorail-Like Gripper System with Dynamic and Modular Reconfiguration for Diverse Finger Layouts β€”
ThI1I.96 DemoDiffusion: One-Shot Human Imitation Using Pre-Trained Diffusion Policy 2506.20668
ThI2I.125 MultiHand: Design and Verification of a Dexterous Hand with Multi-Modal Grasping Capabilities β€”
ThI2I.143 SEM: Enhancing Spatial Understanding for Robust Robot Manipulation 2505.16196
ThI2I.231 Robust Hand Tracking from Visual-Inertial Fusion β€”
ThI2I.235 Adversarial Game-Theoretic Algorithm for Dexterous Grasp Synthesis 2511.05809
ThI2I.26 High-Speed Scooping through Dynamic Manipulation: Model and Practice β€”
ThI2I.292 Generate, Transfer, Adapt: Learning Functional Dexterous Grasping from a Single Human Demonstration 2601.05243
ThI2I.296 TransDexNet: End-to-End Motion Retargeting Network with Transformer for Dexterous Hand Teleoperation from RGB Images β€”
ThI2I.352 Development of the Bioinspired Tendon-Driven DexHand 021 with Proprioceptive Compliance Control 2511.03481
ThI2I.377 Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation 2411.14913
ThI2I.383 DexSinGrasp: Learning a Unified Policy for Dexterous Object Singulation and Grasping in Densely Cluttered Environments 2504.04516
ThI2I.83 FAR-Dex: Few-Shot Data Augmentation and Adaptive Residual Policy Refinement for Dexterous Manipulation 2603.10451
TuAT3.1 Spatially-Anchored Tactile Awareness for Robust Dexterous Manipulation (SaTA) 2510.14647
TuAT3.7 Irrotational Contact Fields 2312.03908
TuAT3.9 Language-Guided Dexterous Functional Grasping by LLM Generated Grasp Functionality and Synergy for Humanoid Manipulation (I) β€”
TuAT4.6 Soft Omni-Functional Robotic Gripper with a Force-Enhanced Pleated Mechanism for High Force and Multi-DoF Manipulation β€”
TuBT3.3 CoorGrasp: Coordinated Contact Control for Adaptive Dexterous Grasping under Uncertainty β€”
TuI1I.124 Dynamic Scoop-and-Flick Manipulation for Rapid Non-Prehensile High-Arc Object Transfer β€”
TuI1I.169 Hydrosoft: Non-Holonomic Hydroelastic Models for Compliant Tactile Manipulation 2509.13126
TuI1I.171 CEDex: Cross-Embodiment Dexterous Grasp Generation at Scale from Human-Like Contact Representations 2509.24661
TuI1I.183 VFP: Variational Flow-Matching Policy for Multi-Modal Robot Manipulation 2508.01622
TuI1I.198 Dexterous Planar Pushing under Uncertain Object Properties: A Contact-Aware Goal-Oriented Approach β€”
TuI1I.199 DexKnot: Generalizable Visuomotor Policy Learning for Dexterous Bag-Knotting Manipulation 2603.07136
TuI1I.206 SoftHand Model-W: A 3D-Printed, Anthropomorphic, Underactuated Robot Hand with Integrated Wrist and Carpal Tunnel 2604.00738
TuI1I.214 Two Degree-of-Freedom Vibratory Transport in a Grasp β€”
TuI1I.249 Suction Leap-Hand: Suction Cups on a Multi-Fingered Hand Enable Embodied Dexterity and In-Hand Teleoperation 2509.20646
TuI1I.292 MiniBEE: A New Form Factor for Compact Bimanual Dexterity 2510.01603
TuI1I.375 Approximating Global Contact-Implicit MPC via Sampling and Local Complementarity 2505.13350
TuI1I.52 A Gripper with Extreme Stiffness Anisotropy for High-Speed Handling of Fragile Foods β€”
TuI1I.67 Design of an Adaptive Modular Anthropomorphic Dexterous Hand for Human-Like Manipulation 2511.22100
TuI1I.79 Multi-Modal Affordance Planner with Temporal-Context Action Policy for Long-Horizon Bimanual Robot Manipulation β€”
TuI1LB.4 A Wire-Driven Robotic Hand with Mode-Switchable Planetary Transmission for Dynamic Manipulation β€”
TuI2I.11 Nonlinear Model Predictive Control for Robotic Pushing of Planar Objects with Generic Shape β€”
TuI2I.111 Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-The-Wild Human Demonstrations (AINA) 2511.16661
TuI2I.141 DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands 2412.01791
TuI2I.169 One-Policy-Fits-All: Geometry-Aware Action Latents for Cross-Embodiment Manipulation 2603.14522
TuI2I.173 Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt 2505.20795
TuI2I.196 UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data 2603.05312
TuI2I.218 Diffusing Trajectory Optimization Problems for Recovery During Multi-Finger Manipulation 2510.07030
TuI2I.309 Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation 2510.08807
TuI2I.337 Spectral Decomposition of Inverse Dynamics for Fast Exploration in Model-Based Manipulation 2603.27796
TuI2I.34 The Folding Hand: Anthropomorphic Robotic Hands with a Compact Reconfigurable Humanoid Palm Design β€”
TuI2I.67 Learning Geometry-Aware Nonprehensile Pushing and Pulling with Dexterous Hands 2509.18455
TuI2I.95 Approximated Collision Detection for Contact-Rich Dexterous Manipulation with Nonnegative Least Squares β€”
WeAT1.2 Robotic Dexterous Manipulation via Anisotropic Friction Modulation Using Passive Rollers 2603.27452
WeAT1.5 Push Anything: Single and Multi-Object Pushing from First Sight with Contact-Implicit MPC 2510.19974
WeI1I.109 Touch-Based Object Localisation with Spatially-Aware Belief Entropy Estimation β€”
WeI1I.165 Bi-Hap: A Bi-Directional Learning-Based Control and Momentum-Based Haptic Feedback System for Dexterous In-Hand Telemanipulation 2409.20527
WeI1I.183 HybNetic: A Mobile Hybrid Magnetic Actuation System β€”
WeI1I.213 DexCtrl: Sim-To-Real Dexterity with Adaptive Controller Learning 2505.00991
WeI1I.345 Frictional and Prismatic Pin-Array Gripper for Universal Gripping and Stable Tool Manipulation β€”
WeI1I.36 A Multi-Level Similarity Approach for Single-View Object Grasping: Matching, Planning, and Fine-Tuning 2507.11938
WeI1I.376 UniFucGrasp: Human-Hand-Inspired Unified Functional Grasp Annotation Strategy and Dataset for Diverse Dexterous Hands 2508.03339
WeI1I.386 On Transient Release Dynamics in Robot Throwing: A Sliding Pivot Model β€”
WeI1I.411 Deformable Cluster Manipulation via Whole-Arm Policy Learning 2507.17085
WeI2I.101 Best of Sim and Real: Decoupled Visuomotor Manipulation via Learning Control in Simulation and Perception in Real 2509.25747
WeI2I.13 Enhancing Safety and Manipulability of Redundant Manipulators: Accelerated Motion Generation in Dynamic Environments β€”
WeI2I.176 T(R,O) Grasp: Efficient Graph Diffusion of Robot-Object Spatial Transformation for Cross-Embodiment Dexterous Grasping 2510.12724
WeI2I.204 Learning Dexterous Manipulation with Quantized Hand State (DQ-RISE) 2509.17450
WeI2I.292 Towards Automated Chicken Deboning via Learning-Based Dynamically-Adaptive 6-DoF Multi-Material Cutting 2510.15376
WeI2I.338 Manipulating Elasto-Plastic Objects with 3D Occupancy and Learning-Based Predictive Control 2505.16249
WeI2I.428 Estimating Deformable-Rigid Contact Interactions for a Deformable Tool via Learning and Model-Based Optimization 2505.10884
WeI2I.55 Flow before Imitation: Learning Dexterous In-Hand Manipulation with Dynamic Visuotactile Shortcut Policy (FBI) 2508.14441
WeI2LB.10 Ultra-Low-Impedance Robotic Gripper for High-Bandwidth and Transparent Physical Interaction β€”

Related

← Back to ICRA-2026-VLA-Manipulation-Survey

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally