Skip to content

ICRA 2026 Topic Grasping

Heungwoo edited this page Jun 1, 2026 · 2 revisions

ICRA 2026 β€” Grasp Synthesis & Grippers (Topic Analysis)

Venue: IEEE ICRA 2026 Β· Vienna, Austria Β· June 1–5, 2026 Compiled against the official PaperCept program. Paper IDs (e.g. ThI1I.233) are session/program codes; arXiv IDs and quantitative numbers below are verified against the public arXiv/project pages where noted, and never fabricated.

This page surveys 41 papers in the grasp-synthesis and end-effector cluster of ICRA 2026. The cluster spans the full pipeline from 6-DoF grasp pose detection on point clouds, through generative grasp synthesis (diffusion / flow-matching), language- and sketch-conditioned grasping, and closed-loop / reactive grasping, down to novel gripper hardware (soft, adhesive, underactuated, multifingered). Two structural shifts stand out versus prior years: (1) diffusion / flow-matching has become the default generative backbone for SE(3) grasp synthesis (GraspGen, DAGDiff, HOGraspFlow, DiffuDepGrasp); and (2) large simulated grasp datasets and grasp "foundation"-style generators trained across embodiments (GraspGen's 53M-grasp dataset; Grasp-MPC's 2M-trajectory value function) are pushing the field toward turnkey, cross-gripper grasping. The hardware half of the cluster is unusually strong this year, with a wave of bio-inspired and food-/fabric-/agriculture-specialized grippers.

For the broader manipulation context see ICRA 2026 Survey; for dexterous/multifinger work see Dexterous Manipulation review.

Sub-trends

1. Generative 6-DoF grasp synthesis (diffusion & flow matching)

The dominant methodological theme. GraspGen (ThI1I.73) is the flagship: a Diffusion-Transformer grasp generator paired with an on-generator-trained discriminator that scores/filters samples, released with a simulated dataset of over 53M grasps across multiple grippers β€” explicitly framed as a turnkey, cross-embodiment 6-DoF grasping foundation. DAGDiff (TuI2I.207) extends diffusion to the dual-arm setting, denoising directly in SE(3)Γ—SE(3) and steering generation with geometry-, stability-, and collision-aware classifier guidance rather than region heuristics. HOGraspFlow (ThI1I.70) uses denoising flow matching to retarget a single RGB hand-object-interaction image into multi-modal parallel-jaw grasps, conditioned on RGB foundation features, HOI contact reconstruction, and a taxonomy-aware grasp-type prior. DiffuDepGrasp (TuI1I.328) applies diffusion in a different place β€” modeling realistic depth-sensor noise to close the sim-to-real gap for depth-based grasp policies. A Hybrid Optimization Framework for Grasp Synthesis under Partial Observations (ThI2I.96) blends a learned energy-based model with analytical ICP inside a Stein variational scheme for grasps from partial point clouds.

2. Language / semantic / instruction-conditioned grasping

Grasping is increasingly conditioned on natural-language and multimodal instructions. GraspControl (ThI1I.374) introduces a text-plus-sketch interface: it augments language with grasp position/orientation and gripper sketches, then generates 2D grasp sketches for controllable synthesis. GarmentPile++ (ThI1I.329) couples VLM vision-language reasoning with a learned retrieval-affordance model to decide which / where / how to retrieve a single garment from a pile, including when dual-arm cooperation is needed. Efficient Alignment of Unconditioned Action Prior for Language-Conditioned Pick and Place in Clutter (ThI2I.389) aligns an unconditioned action prior to language goals for target grasping in open clutter. LACY (WeI1I.298) builds a bidirectional language-action cycle (L2A / A2L / consistency) inside one VLM, enabling self-improving manipulation data generation. Tool-Grasp (WeI1I.150) targets functional 6-DoF grasps for general-purpose hand tools, addressing the scarcity of fine-grained functional-grasp annotations.

3. Dexterous, multifinger & hand-retargeting grasping

Beyond parallel jaws. HOGraspFlow (ThI1I.70, above) is taxonomy-aware over human grasp types. Pose Retargeting from a Single RGB Camera (TuI2I.86) does optimization-based hand-pose and wrist-pose estimation for vision-based teleoperation data collection β€” a key enabler for dexterous imitation. Taxonomy-Aware Dynamic Motion Generation on Hyperbolic Manifolds (TuI1I.113) embeds biomechanical grasp/motion taxonomies in hyperbolic space for human-like motion. On the hardware side, A Cable-Driven Soft Robotic Hand with an In-Hand RGB-D Camera (WeI1I.433) gives each finger independent actuation plus integrated in-hand vision for dexterous grasping, and the reconfigurable multifingered gripper for agri-food (WeI2LB.13) studies topology-dependent interaction patterns for dexterous food handling.

4. Grasp foundation models, large datasets & benchmarking

A clear push toward scale and reproducible evaluation. GraspGen's 53M-grasp dataset and Grasp-MPC's value function trained on 2M grasp trajectories (WeI2I.95) are the two largest-scale efforts. Benchmarking is a sub-cluster in its own right: A Benchmarking Study of Vision-Based Robotic Grasping Algorithms (TuI2I.29) compares learning-based vs analytical methods under a shared protocol, and Benchmarking the Effects of Object Pose Estimation and Reconstruction on Robotic Grasping Success (TuI2I.205) isolates how upstream perception errors propagate to grasp success. GSWorld (TuI1I.48) provides a closed-loop, photo-realistic 3D-Gaussian-Splatting + physics simulator (GSDF asset format, 3 embodiments, 40+ objects) for reproducible policy evaluation and sim-to-real. Communication-Efficient Module-Wise Federated Learning for Grasp Pose Detection (WeI2I.366) tackles the data-privacy/centralization problem of training GPD on large datasets.

5. Closed-loop, reactive & active-perception grasping

Moving past open-loop "predict-then-execute." Grasp-MPC (WeI2I.95) wraps a value function in model-predictive control for closed-loop, reactive 6-DoF grasping of novel objects in clutter. Hierarchical Reactive Grasping (TuI1I.142) combines task-space velocity fields with joint-space QP for fast reactive, collision-free grasping on high-DoF arms. Active perception recurs: HEAPGrasp (TuI1I.425) does hand-eye active perception to grasp objects of diverse optical properties; GPD-AP (WeI2I.126) drives viewpoint selection from grasp-pose feedback for occlusion-robust manipulation; and Active Perception for Deformable Linear Objects Stiffness Estimation (TuI2LB.17) actively probes DLOs to infer hidden material properties. Tracing Energy Flow (TuI2I.403) learns tactile grasping-force control to reduce slippage under dynamic, multi-contact interaction.

6. Grasping in clutter, decluttering & pick-and-place planning

Clutter is the recurring hard setting. GAPG (WeI1I.305) learns a geometry-aware push-grasping synergy for goal-oriented manipulation, using pushing to create graspable configurations. Leveraging Embodied Mechanical Intelligence for Learning Decluttering Tasks (TuAT3.8) studies how a DRL grasp planner behaves with a soft-rigid "Soft ScoopGripper." On the planning side: Differentiable Optimization-Based Modular Planning for Pick-and-Place with Regrasp (ThI2I.124) replaces sampling with differentiable optimization for repeated regrasp; Learning from Planned Data to Improve Pick-and-Place Planning Efficiency (WeI2I.362) predicts shared grasps feasible at both pick and place; and RoboPacker (WeI2I.390) is a full autonomous packing system using open-vocabulary shape completion + hierarchical RL for dense box packing.

7. Novel gripper & end-effector hardware

The largest hardware contingent in years, much of it application-specialized. Soft/adhesive: A Dual-Adhesion-Enhanced Soft Gripper with Microwedge Adhesives and SMA-Driven Microspines (WeI2I.22, lizard-inspired dual adhesion for smooth and rough surfaces); Structural Interlocking-Based Weaving Gripper (ThI1LB.2); A Tactile Rubbing Gripper for Reliable Fabric Separation (ThI1I.300); A Novel Soft Gripper with a Unilateral Fingernail-Like Mechanism for grasping flat objects (TuI1I.283); Soft Self-Centering Gripper for Delicate Object Handling (ThI2I.211, fruit/vegetable). Food / agriculture / industry: An Underactuated Robotic Gripper with Flowability and Variable Stiffness for Food Bin-Picking (WeI2LB.8); Curvature-Adaptable Robotic End-Effectors (WeI2LB.17, automotive assembly of curved parts); reconfigurable multifingered agri-food gripper (WeI2LB.13). Field/specialized: SureGrip (WeI2I.82) detects natural handholds and evaluates grasp quality for free-climbing robots on rock/lunar-cave terrain. Sensing-integrated systems: A Hyperspectral Imaging Guided Robotic Grasping System (TuI1I.9); the in-hand-camera soft hand (WeI1I.433); Zero-Shot Exocentric Viewpoint-Robust Imitation Learning (TuI1I.117, handheld-gripper data collection); GIFT (ThI1I.233, geometry-induced functional skill transfer); Zero-Shot Recognition of Test Tube Types (WeI1I.366, life-science automation).

Standout deep-dives

  • GraspGen (ThI1I.73) β€” arXiv 2507.13097 (NVIDIA et al.). A DiffusionTransformer grasp generator + on-generator-trained discriminator, released with a simulated dataset of 53M+ grasps across grippers; reaches state-of-the-art on the FetchBench grasping benchmark and is positioned as a turnkey cross-embodiment 6-DoF grasping framework. The clearest "grasp foundation model" entry in the cluster.

  • DAGDiff (TuI2I.207) β€” arXiv 2509.21145 (ICRA 2026; code at github.com/DAG-Diff/dual-arm-grasp-diffusion). End-to-end diffusion that denoises dual-arm grasp pairs directly in SE(3)Γ—SE(3), guided by geometry-, stability-, and collision-aware terms toward force-closure-compliant grasps; validated with analytical force-closure checks, collision analysis, large-scale physics sim, and real heterogeneous dual-arm execution on unseen objects.

  • HOGraspFlow (ThI1I.70) β€” arXiv 2509.16871 (KIT). Flow-matching SE(3) grasp synthesis that retargets a single RGB hand-object-interaction image into multi-modal parallel-jaw grasps with no explicit object geometry prior, conditioned on RGB foundation features, HOI contact reconstruction, and a taxonomy-aware grasp-type prior; reports an average real-world success rate above 83%.

  • DiffuDepGrasp (TuI1I.328) β€” arXiv 2511.12912. A sim-to-real framework whose Diffusion Depth Generator synthesizes sensor-realistic depth noise (Diffusion Depth Module + Noise Grafting Module), enabling zero-shot transfer of depth-based grasp policies with a reported 95.7% average success rate on 12-object grasping and strong generalization to unseen objects β€” using only raw depth at deployment.

  • Grasp-MPC (WeI2I.95) β€” arXiv 2509.06201 (Oxford/NVIDIA). A closed-loop 6-DoF grasping policy: a value function trained on a synthetic dataset of 2M grasp trajectories (successes and failures) deployed inside an MPC framework with collision-avoidance and smoothness cost terms, for robust reactive grasping of novel objects in clutter.

  • GarmentPile++ (ThI1I.329) β€” arXiv 2603.04158. Affordance-driven cluttered-garment retrieval that integrates VLM vision-language reasoning with a Retrieval Affordance Model across three stages β€” which to retrieve (SAM2 masks + VLM), where (affordance grasp points), and how (VLM decides single- vs dual-arm) β€” guaranteeing exactly one garment per attempt.

Complete paper list (41)

Code Title arXiv
ThI1I.233 GIFT: Geometry-Induced Functional Transfer for Category-Level Object Manipulation 2503.15371
ThI1I.300 A Tactile Rubbing Gripper for Reliable Fabric Separation β€”
ThI1I.329 GarmentPile++: Affordance-Driven Cluttered Garments Retrieval with Vision-Language Reasoning 2603.04158
ThI1I.374 GraspControl: Text-Sketch Instruction As an Interface for Controllable Grasp Synthesis β€”
ThI1I.70 HOGraspFlow: Taxonomy-Aware Hand-Object Retargeting for Multi-Modal SE(3) Grasp Generation 2509.16871
ThI1I.73 GraspGen: A Diffusion-Based Framework for 6-DOF Grasping with On-Generator Training 2507.13097
ThI1LB.2 Structural Interlocking-Based Weaving Gripper for Enhanced Grasping Performance β€”
ThI2I.124 Differentiable Optimization-Based Modular Planning Framework for Pick-And-Place with Regrasp β€”
ThI2I.211 Design and Validation of a Soft Self-Centering Gripper for Delicate Object Handling β€”
ThI2I.389 Efficient Alignment of Unconditioned Action Prior for Language-Conditioned Pick and Place in Clutter (I) 2503.09423
ThI2I.96 A Hybrid Optimization Framework for Grasp Synthesis under Partial Observations β€”
TuAT3.8 Leveraging Embodied Mechanical Intelligence for Learning Decluttering Tasks β€”
TuI1I.113 Taxonomy-Aware Dynamic Motion Generation on Hyperbolic Manifolds 2509.21281
TuI1I.117 Zero-Shot Exocentric Viewpoint-Robust Imitation Learning (VIL): Bridging Handheld Gripper and Exocentric Views β€”
TuI1I.142 Hierarchical Reactive Grasping Via Task-Space Velocity Fields and Joint-Space Quadratic Programming 2509.01044
TuI1I.283 A Novel Soft Gripper Design Integrating a Unilateral Fingernail-Like Mechanism for Grasping Flat Object β€”
TuI1I.328 DiffuDepGrasp: Diffusion-Based Depth Noise Modeling Empowers Sim-To-Real Robotic Grasping 2511.12912
TuI1I.425 HEAPGrasp: Hand-Eye Active Perception to Grasp Objects with Diverse Optical Properties β€”
TuI1I.48 GSWorld: Closed-Loop Photo-Realistic Simulation Suite for Robotic Manipulation 2510.20813
TuI1I.9 A Hyperspectral Imaging Guided Robotic Grasping System 2512.05578
TuI2I.205 Benchmarking the Effects of Object Pose Estimation and Reconstruction on Robotic Grasping Success 2602.17101
TuI2I.207 DAGDiff: Guiding Dual-Arm Grasp Diffusion to Stable and Collision-Free Grasps 2509.21145
TuI2I.29 A Benchmarking Study of Vision-Based Robotic Grasping Algorithms 2503.11163
TuI2I.403 Tracing Energy Flow: Learning Tactile-Based Grasping Force Control to Reduce Slippage in Dynamic Object Interaction 2512.21043
TuI2I.86 Pose Retargeting from a Single RGB Camera: Optimization-Based Hand Pose Retargeting and Wrist Pose Estimation β€”
TuI2LB.17 Active Perception for Deformable Linear Objects Stiffness Estimation β€”
WeI1I.150 Tool-Grasp: A 6-DoF Functional Grasping Framework for General-Purpose Hand Tools β€”
WeI1I.298 LACY: A Vision-Language Model-Based Language-Action Cycle for Self-Improving Robotic Manipulation 2511.02239
WeI1I.305 GAPG: Geometry Aware Push-Grasping Synergy for Goal-Oriented Manipulation in Clutter 2603.21195
WeI1I.366 Zero-Shot Recognition of Test Tube Types by Automatically Collecting and Labeling RGB Data β€”
WeI1I.433 A Cable-Driven Soft Robotic Hand with an In-Hand RGB-D Camera for Dexterous Grasping and Manipulation β€”
WeI2I.126 GPD-AP: A Grasp Pose-Driven Active Perception Framework for Occlusion-Robust Robotic Manipulation β€”
WeI2I.22 A Dual-Adhesion-Enhanced Soft Gripper with Microwedge Adhesives and SMA-Driven Microspines β€”
WeI2I.362 Learning from Planned Data to Improve Robotic Pick-And-Place Planning Efficiency 2506.15920
WeI2I.366 Communication-Efficient Module-Wise Federated Learning for Grasp Pose Detection in Cluttered Environments 2507.05861
WeI2I.390 RoboPacker: An Autonomous Robotic Packing System for General Objects (I) β€”
WeI2I.82 SureGrip: Perceptual Grasping of Natural Handholds for Free-Climbing Robots β€”
WeI2I.95 Grasp-MPC: Closed-Loop Visual Grasping Via Value-Guided Model Predictive Control 2509.06201
WeI2LB.13 Towards Dexterous Agri-Food Manipulation: Topology-Dependent Interaction Patterns in a Reconfigurable Multifingered Gripper β€”
WeI2LB.17 Curvature Adaptable Robotic End-Effectors β€”
WeI2LB.8 An Underactuated Robotic Gripper with Flowability and Variable Stiffness for Food Bin-Picking β€”

Related

← Back to ICRA-2026-VLA-Manipulation-Survey

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally