-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 DexEvolve
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: Manipulation 2 Β· paper #59 Authors: RenΓ© ZurbrΓΌgg, Andrei Cramariuc, Marco Hutter arXiv: 2602.15201 Β· program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1: diverse, physically stable XHand grasps produced by evolutionary refinement, shown on two Handles assets (pink handles on the tabletop) and one object asset β candidates are refined directly inside high-fidelity simulation rather than merely filtered by it.
Data-driven dexterous grasp prediction needs large, diverse datasets, but analytical grasp synthesis makes simplifying assumptions (coarse contact dynamics, friction approximations), so most proposals fail high-fidelity physics verification and get discarded β a sample-inefficient generate-then-filter paradigm. High-fidelity simulators are also non-differentiable, ruling out gradient-based refinement of full grasp configurations.
DexEvolve reinterprets the simulator as a black-box objective: analytical seeds from GraspQP initialize an asynchronous, gradient-free evolutionary algorithm running in Isaac Sim with massively parallel rollouts and early rejection. Grasps G = (wrist pose, joint states, delta joint commands) evolve via density-aware tournament selection (suppressing clustered modes), finger/pose-swapping crossover, Gaussian mutation, and archive-based novelty insertion (candidates within distance Ο of a neighbor only replace it on fitness improvement). Fitness combines a disturbance-protocol lifetime score (forces along Β±x, Β±y, Β±z), a contact-distance penalty, and a penetration penalty; contact points and grasp commands are resampled per offspring via FPS + contact-Jacobian solves. Because the objective need not be differentiable, a PointNet++ preference model trained on ~1,000 human pairwise annotations (BradleyβTerry loss) can steer refinement toward natural grasps. The refined distribution is finally distilled into a point-cloud-conditioned diffusion model (DexGraspAnything-style with contact-consistency constraints) for deployment. The paper also introduces a Handles dataset of 90 geometrically distinct, commercially modeled handle/knob assets annotated for the XHand.
On the Handles dataset and a DexGraspNet subset, refinement yields over 120 distinct stable grasps per object β a 1.7β6x improvement over unrefined analytical seeds β while raising success rates and entropy; convergence plateaus by roughly 10k simulator steps. Against diffusion trained on the same seeds, evolutionary refinement achieves ~115 vs 72 unique grasps at 32 seeds (60% better) and 118 vs 81 at 128 seeds (46% better), with higher entropy (~3.0 vs 2.4β2.6). Real-world deployment on a Franka Panda + XHand (RealSense L415-captured point clouds lifted with Depth Anything V3, cuRobo planning) shows successful cabinet-handle grasps across multiple grasp modes.
Turns high-fidelity simulation from a reject filter into the optimizer itself, and shows quality-diversity machinery (archives, density-aware selection) scaling to full dexterous grasp configurations β a data-generation recipe upstream of the learned-grasping thread in Review-Dexterous-Manipulation.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)