-
Notifications
You must be signed in to change notification settings - Fork 0
ICRA 2026 Topic Grasping
Venue: IEEE ICRA 2026 Β· Vienna, Austria Β· June 1β5, 2026 Compiled against the official PaperCept program. Paper IDs (e.g.
ThI1I.233) are session/program codes; arXiv IDs and quantitative numbers below are verified against the public arXiv/project pages where noted, and never fabricated.
This page surveys 41 papers in the grasp-synthesis and end-effector cluster of ICRA 2026. The cluster spans the full pipeline from 6-DoF grasp pose detection on point clouds, through generative grasp synthesis (diffusion / flow-matching), language- and sketch-conditioned grasping, and closed-loop / reactive grasping, down to novel gripper hardware (soft, adhesive, underactuated, multifingered). Two structural shifts stand out versus prior years: (1) diffusion / flow-matching has become the default generative backbone for SE(3) grasp synthesis (GraspGen, DAGDiff, HOGraspFlow, DiffuDepGrasp); and (2) large simulated grasp datasets and grasp "foundation"-style generators trained across embodiments (GraspGen's 53M-grasp dataset; Grasp-MPC's 2M-trajectory value function) are pushing the field toward turnkey, cross-gripper grasping. The hardware half of the cluster is unusually strong this year, with a wave of bio-inspired and food-/fabric-/agriculture-specialized grippers.
For the broader manipulation context see ICRA 2026 Survey; for dexterous/multifinger work see Dexterous Manipulation review.
The dominant methodological theme. GraspGen (ThI1I.73) is the flagship: a Diffusion-Transformer grasp generator paired with an on-generator-trained discriminator that scores/filters samples, released with a simulated dataset of over 53M grasps across multiple grippers β explicitly framed as a turnkey, cross-embodiment 6-DoF grasping foundation. DAGDiff (TuI2I.207) extends diffusion to the dual-arm setting, denoising directly in SE(3)ΓSE(3) and steering generation with geometry-, stability-, and collision-aware classifier guidance rather than region heuristics. HOGraspFlow (ThI1I.70) uses denoising flow matching to retarget a single RGB hand-object-interaction image into multi-modal parallel-jaw grasps, conditioned on RGB foundation features, HOI contact reconstruction, and a taxonomy-aware grasp-type prior. DiffuDepGrasp (TuI1I.328) applies diffusion in a different place β modeling realistic depth-sensor noise to close the sim-to-real gap for depth-based grasp policies. A Hybrid Optimization Framework for Grasp Synthesis under Partial Observations (ThI2I.96) blends a learned energy-based model with analytical ICP inside a Stein variational scheme for grasps from partial point clouds.
Grasping is increasingly conditioned on natural-language and multimodal instructions. GraspControl (ThI1I.374) introduces a text-plus-sketch interface: it augments language with grasp position/orientation and gripper sketches, then generates 2D grasp sketches for controllable synthesis. GarmentPile++ (ThI1I.329) couples VLM vision-language reasoning with a learned retrieval-affordance model to decide which / where / how to retrieve a single garment from a pile, including when dual-arm cooperation is needed. Efficient Alignment of Unconditioned Action Prior for Language-Conditioned Pick and Place in Clutter (ThI2I.389) aligns an unconditioned action prior to language goals for target grasping in open clutter. LACY (WeI1I.298) builds a bidirectional language-action cycle (L2A / A2L / consistency) inside one VLM, enabling self-improving manipulation data generation. Tool-Grasp (WeI1I.150) targets functional 6-DoF grasps for general-purpose hand tools, addressing the scarcity of fine-grained functional-grasp annotations.
Beyond parallel jaws. HOGraspFlow (ThI1I.70, above) is taxonomy-aware over human grasp types. Pose Retargeting from a Single RGB Camera (TuI2I.86) does optimization-based hand-pose and wrist-pose estimation for vision-based teleoperation data collection β a key enabler for dexterous imitation. Taxonomy-Aware Dynamic Motion Generation on Hyperbolic Manifolds (TuI1I.113) embeds biomechanical grasp/motion taxonomies in hyperbolic space for human-like motion. On the hardware side, A Cable-Driven Soft Robotic Hand with an In-Hand RGB-D Camera (WeI1I.433) gives each finger independent actuation plus integrated in-hand vision for dexterous grasping, and the reconfigurable multifingered gripper for agri-food (WeI2LB.13) studies topology-dependent interaction patterns for dexterous food handling.
A clear push toward scale and reproducible evaluation. GraspGen's 53M-grasp dataset and Grasp-MPC's value function trained on 2M grasp trajectories (WeI2I.95) are the two largest-scale efforts. Benchmarking is a sub-cluster in its own right: A Benchmarking Study of Vision-Based Robotic Grasping Algorithms (TuI2I.29) compares learning-based vs analytical methods under a shared protocol, and Benchmarking the Effects of Object Pose Estimation and Reconstruction on Robotic Grasping Success (TuI2I.205) isolates how upstream perception errors propagate to grasp success. GSWorld (TuI1I.48) provides a closed-loop, photo-realistic 3D-Gaussian-Splatting + physics simulator (GSDF asset format, 3 embodiments, 40+ objects) for reproducible policy evaluation and sim-to-real. Communication-Efficient Module-Wise Federated Learning for Grasp Pose Detection (WeI2I.366) tackles the data-privacy/centralization problem of training GPD on large datasets.
Moving past open-loop "predict-then-execute." Grasp-MPC (WeI2I.95) wraps a value function in model-predictive control for closed-loop, reactive 6-DoF grasping of novel objects in clutter. Hierarchical Reactive Grasping (TuI1I.142) combines task-space velocity fields with joint-space QP for fast reactive, collision-free grasping on high-DoF arms. Active perception recurs: HEAPGrasp (TuI1I.425) does hand-eye active perception to grasp objects of diverse optical properties; GPD-AP (WeI2I.126) drives viewpoint selection from grasp-pose feedback for occlusion-robust manipulation; and Active Perception for Deformable Linear Objects Stiffness Estimation (TuI2LB.17) actively probes DLOs to infer hidden material properties. Tracing Energy Flow (TuI2I.403) learns tactile grasping-force control to reduce slippage under dynamic, multi-contact interaction.
Clutter is the recurring hard setting. GAPG (WeI1I.305) learns a geometry-aware push-grasping synergy for goal-oriented manipulation, using pushing to create graspable configurations. Leveraging Embodied Mechanical Intelligence for Learning Decluttering Tasks (TuAT3.8) studies how a DRL grasp planner behaves with a soft-rigid "Soft ScoopGripper." On the planning side: Differentiable Optimization-Based Modular Planning for Pick-and-Place with Regrasp (ThI2I.124) replaces sampling with differentiable optimization for repeated regrasp; Learning from Planned Data to Improve Pick-and-Place Planning Efficiency (WeI2I.362) predicts shared grasps feasible at both pick and place; and RoboPacker (WeI2I.390) is a full autonomous packing system using open-vocabulary shape completion + hierarchical RL for dense box packing.
The largest hardware contingent in years, much of it application-specialized. Soft/adhesive: A Dual-Adhesion-Enhanced Soft Gripper with Microwedge Adhesives and SMA-Driven Microspines (WeI2I.22, lizard-inspired dual adhesion for smooth and rough surfaces); Structural Interlocking-Based Weaving Gripper (ThI1LB.2); A Tactile Rubbing Gripper for Reliable Fabric Separation (ThI1I.300); A Novel Soft Gripper with a Unilateral Fingernail-Like Mechanism for grasping flat objects (TuI1I.283); Soft Self-Centering Gripper for Delicate Object Handling (ThI2I.211, fruit/vegetable). Food / agriculture / industry: An Underactuated Robotic Gripper with Flowability and Variable Stiffness for Food Bin-Picking (WeI2LB.8); Curvature-Adaptable Robotic End-Effectors (WeI2LB.17, automotive assembly of curved parts); reconfigurable multifingered agri-food gripper (WeI2LB.13). Field/specialized: SureGrip (WeI2I.82) detects natural handholds and evaluates grasp quality for free-climbing robots on rock/lunar-cave terrain. Sensing-integrated systems: A Hyperspectral Imaging Guided Robotic Grasping System (TuI1I.9); the in-hand-camera soft hand (WeI1I.433); Zero-Shot Exocentric Viewpoint-Robust Imitation Learning (TuI1I.117, handheld-gripper data collection); GIFT (ThI1I.233, geometry-induced functional skill transfer); Zero-Shot Recognition of Test Tube Types (WeI1I.366, life-science automation).
-
GraspGen (
ThI1I.73) β arXiv 2507.13097 (NVIDIA et al.). A DiffusionTransformer grasp generator + on-generator-trained discriminator, released with a simulated dataset of 53M+ grasps across grippers; reaches state-of-the-art on the FetchBench grasping benchmark and is positioned as a turnkey cross-embodiment 6-DoF grasping framework. The clearest "grasp foundation model" entry in the cluster. -
DAGDiff (
TuI2I.207) β arXiv 2509.21145 (ICRA 2026; code at github.com/DAG-Diff/dual-arm-grasp-diffusion). End-to-end diffusion that denoises dual-arm grasp pairs directly in SE(3)ΓSE(3), guided by geometry-, stability-, and collision-aware terms toward force-closure-compliant grasps; validated with analytical force-closure checks, collision analysis, large-scale physics sim, and real heterogeneous dual-arm execution on unseen objects. -
HOGraspFlow (
ThI1I.70) β arXiv 2509.16871 (KIT). Flow-matching SE(3) grasp synthesis that retargets a single RGB hand-object-interaction image into multi-modal parallel-jaw grasps with no explicit object geometry prior, conditioned on RGB foundation features, HOI contact reconstruction, and a taxonomy-aware grasp-type prior; reports an average real-world success rate above 83%. -
DiffuDepGrasp (
TuI1I.328) β arXiv 2511.12912. A sim-to-real framework whose Diffusion Depth Generator synthesizes sensor-realistic depth noise (Diffusion Depth Module + Noise Grafting Module), enabling zero-shot transfer of depth-based grasp policies with a reported 95.7% average success rate on 12-object grasping and strong generalization to unseen objects β using only raw depth at deployment. -
Grasp-MPC (
WeI2I.95) β arXiv 2509.06201 (Oxford/NVIDIA). A closed-loop 6-DoF grasping policy: a value function trained on a synthetic dataset of 2M grasp trajectories (successes and failures) deployed inside an MPC framework with collision-avoidance and smoothness cost terms, for robust reactive grasping of novel objects in clutter. -
GarmentPile++ (
ThI1I.329) β arXiv 2603.04158. Affordance-driven cluttered-garment retrieval that integrates VLM vision-language reasoning with a Retrieval Affordance Model across three stages β which to retrieve (SAM2 masks + VLM), where (affordance grasp points), and how (VLM decides single- vs dual-arm) β guaranteeing exactly one garment per attempt.
| Code | Title | arXiv |
|---|---|---|
| ThI1I.233 | GIFT: Geometry-Induced Functional Transfer for Category-Level Object Manipulation | 2503.15371 |
| ThI1I.300 | A Tactile Rubbing Gripper for Reliable Fabric Separation | β |
| ThI1I.329 | GarmentPile++: Affordance-Driven Cluttered Garments Retrieval with Vision-Language Reasoning | 2603.04158 |
| ThI1I.374 | GraspControl: Text-Sketch Instruction As an Interface for Controllable Grasp Synthesis | β |
| ThI1I.70 | HOGraspFlow: Taxonomy-Aware Hand-Object Retargeting for Multi-Modal SE(3) Grasp Generation | 2509.16871 |
| ThI1I.73 | GraspGen: A Diffusion-Based Framework for 6-DOF Grasping with On-Generator Training | 2507.13097 |
| ThI1LB.2 | Structural Interlocking-Based Weaving Gripper for Enhanced Grasping Performance | β |
| ThI2I.124 | Differentiable Optimization-Based Modular Planning Framework for Pick-And-Place with Regrasp | β |
| ThI2I.211 | Design and Validation of a Soft Self-Centering Gripper for Delicate Object Handling | β |
| ThI2I.389 | Efficient Alignment of Unconditioned Action Prior for Language-Conditioned Pick and Place in Clutter (I) | 2503.09423 |
| ThI2I.96 | A Hybrid Optimization Framework for Grasp Synthesis under Partial Observations | β |
| TuAT3.8 | Leveraging Embodied Mechanical Intelligence for Learning Decluttering Tasks | β |
| TuI1I.113 | Taxonomy-Aware Dynamic Motion Generation on Hyperbolic Manifolds | 2509.21281 |
| TuI1I.117 | Zero-Shot Exocentric Viewpoint-Robust Imitation Learning (VIL): Bridging Handheld Gripper and Exocentric Views | β |
| TuI1I.142 | Hierarchical Reactive Grasping Via Task-Space Velocity Fields and Joint-Space Quadratic Programming | 2509.01044 |
| TuI1I.283 | A Novel Soft Gripper Design Integrating a Unilateral Fingernail-Like Mechanism for Grasping Flat Object | β |
| TuI1I.328 | DiffuDepGrasp: Diffusion-Based Depth Noise Modeling Empowers Sim-To-Real Robotic Grasping | 2511.12912 |
| TuI1I.425 | HEAPGrasp: Hand-Eye Active Perception to Grasp Objects with Diverse Optical Properties | β |
| TuI1I.48 | GSWorld: Closed-Loop Photo-Realistic Simulation Suite for Robotic Manipulation | 2510.20813 |
| TuI1I.9 | A Hyperspectral Imaging Guided Robotic Grasping System | 2512.05578 |
| TuI2I.205 | Benchmarking the Effects of Object Pose Estimation and Reconstruction on Robotic Grasping Success | 2602.17101 |
| TuI2I.207 | DAGDiff: Guiding Dual-Arm Grasp Diffusion to Stable and Collision-Free Grasps | 2509.21145 |
| TuI2I.29 | A Benchmarking Study of Vision-Based Robotic Grasping Algorithms | 2503.11163 |
| TuI2I.403 | Tracing Energy Flow: Learning Tactile-Based Grasping Force Control to Reduce Slippage in Dynamic Object Interaction | 2512.21043 |
| TuI2I.86 | Pose Retargeting from a Single RGB Camera: Optimization-Based Hand Pose Retargeting and Wrist Pose Estimation | β |
| TuI2LB.17 | Active Perception for Deformable Linear Objects Stiffness Estimation | β |
| WeI1I.150 | Tool-Grasp: A 6-DoF Functional Grasping Framework for General-Purpose Hand Tools | β |
| WeI1I.298 | LACY: A Vision-Language Model-Based Language-Action Cycle for Self-Improving Robotic Manipulation | 2511.02239 |
| WeI1I.305 | GAPG: Geometry Aware Push-Grasping Synergy for Goal-Oriented Manipulation in Clutter | 2603.21195 |
| WeI1I.366 | Zero-Shot Recognition of Test Tube Types by Automatically Collecting and Labeling RGB Data | β |
| WeI1I.433 | A Cable-Driven Soft Robotic Hand with an In-Hand RGB-D Camera for Dexterous Grasping and Manipulation | β |
| WeI2I.126 | GPD-AP: A Grasp Pose-Driven Active Perception Framework for Occlusion-Robust Robotic Manipulation | β |
| WeI2I.22 | A Dual-Adhesion-Enhanced Soft Gripper with Microwedge Adhesives and SMA-Driven Microspines | β |
| WeI2I.362 | Learning from Planned Data to Improve Robotic Pick-And-Place Planning Efficiency | 2506.15920 |
| WeI2I.366 | Communication-Efficient Module-Wise Federated Learning for Grasp Pose Detection in Cluttered Environments | 2507.05861 |
| WeI2I.390 | RoboPacker: An Autonomous Robotic Packing System for General Objects (I) | β |
| WeI2I.82 | SureGrip: Perceptual Grasping of Natural Handholds for Free-Climbing Robots | β |
| WeI2I.95 | Grasp-MPC: Closed-Loop Visual Grasping Via Value-Guided Model Predictive Control | 2509.06201 |
| WeI2LB.13 | Towards Dexterous Agri-Food Manipulation: Topology-Dependent Interaction Patterns in a Reconfigurable Multifingered Gripper | β |
| WeI2LB.17 | Curvature Adaptable Robotic End-Effectors | β |
| WeI2LB.8 | An Underactuated Robotic Gripper with Flowability and Variable Stiffness for Food Bin-Picking | β |
β Back to ICRA-2026-VLA-Manipulation-Survey
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)