-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 DexMove
Venue: ICLR 2026 Affiliation: ShanghaiTech Β· BIGAI Β· Beihang Category: Dexterous Manipulation β Non-Prehensile / Tactile / Wrist-Finger Synergy Trend tag: Tactile Β· contact-rich Β· flow-matching policy Β· wearable data
flowchart LR
subgraph SIM[Simulation pipeline β 2M sequences]
YCB[88 YCB objects β 352 instances<br/>random scale + rotation] --> Contact[Contact establishment<br/>uniform wrist pose sample + IK]
Contact --> Filter[Reject self-collision / penetration<br/>412k valid configurations]
Filter --> Reject[Rejection sampling:<br/>fingertips stable contact over 50 cm displacement]
Reject --> Force[Synthesize force-conditioned trajectories<br/>indentation depth as force surrogate<br/>Gaussian augmentation along contact normal]
end
subgraph REAL[Wearable tactile capture]
Wear[R-Tac fingertip sensors<br/>OV9281 mono camera 120 FPS<br/>33 markers per finger] --> H[20 objects, ~300k frames @ 30 FPS]
end
SIM --> EC[Establish-Contact FM policy<br/>PointNet++ + FiLM + flow matching]
H --> TaFo[TaFo-Net<br/>per-finger spatial enc β cross-finger attn β<br/>finger-wise causal temporal attn]
SIM --> POL[DexMove-Policy<br/>Transformer enc+dec flow matching<br/>past Tp=5, future Tf=5, 30 Hz]
TaFo --> POL
POL --> Run[Franka FR3 + Allegro Hand<br/>~22 ms inference / chunk]
Non-prehensile manipulation (pushing, sliding, pivoting without enclosure grasps) with multi-fingered dexterous hands is largely unexplored. Two specific blockers:
- Data scarcity. No large-scale dataset covers force-aware, multi-finger non-prehensile trajectories. Teleoperation suffers from missing haptic feedback (lower fidelity, lower success rates); pure simulation has soft-body and contact-modeling gaps; tactile gloves have layout mismatches between human hands and robot hands.
- No wrist-finger coordination policy. Existing dexterous manipulation focuses on grasping; pushing/pivoting work uses single-contact tools or grippers. Multi-fingered hands couple wrist and finger forces through hand-object dynamics, but no planner exists for this combined control problem.
The paper's argument: multi-fingered hands are intrinsically better for non-prehensile work because they can establish distributed contacts that are more stable than a single-contact rod or two-finger gripper β particularly for thin, cylindrical, or round objects where pushing dynamics are otherwise unpredictable.
Hand-object contact establishment. Uniformly sample wrist poses (Rβ^wrist, Tβ^wrist). For each fingertip, compute displacement d to the nearest object surface; perturb with Gaussian noise Ξ΅ to produce dΜ = d + Ξ΅. Solve IK with:
Γβ^hand = argmin βFK(Aβ^hand, Rβ^wrist, Tβ^wrist) β Pβ^TIPβΒ² + w_pinch Β· L_region
where L_region = βd^TIP β dΜβΒ² encourages contacts inside the tactile sensor's effective region. Using 88 YCB objects Γ random scale & rotation = 352 instances; 1024-2048 candidates per instance β 412k valid contact configurations after collision filtering.
Force-conditioned trajectories. Repositioning controlled by (A^hand, R^wrist, T^wrist) inducing 3-DoF object motion (x, y, yaw). Instead of iLQR (Li & Todorov 2004), the authors use rejection sampling: in MuJoCo, translate the hand along random directions; accept if all fingertips maintain stable contact over 50 cm displacement.
For each accepted direction, uniformly sample object target (P_target^obj, Ο_target^obj). Under the no-slip assumption:
P_t^tip = P_t^obj + R_z(Ο_t^obj) (P_0^TIP β P_0^obj), for t = 0, ..., T.
Contact-force synthesis. Approximate normal force from indentation depth:
G β D_sensor = r β distance(P_t^TIP, surface)
Augment by displacing each fingertip along the contact normal: PΜ_t^TIP = P_t^TIP + n Β· N(0, Ο). Recover joint and wrist configs via IK with a wrist-motion regulariser L^wrist so that the solution biases toward finger-driven rather than arm-driven manipulation. Filter trajectories that leave the workspace. Total: 2M sequences.
A wearable exoskeleton, isomorphic to the Allegro Hand, mounts R-Tac vision-based tactile sensors (Lin et al. 2025) on each human fingertip:
- Camera: OV9281 global-shutter monochrome, 160Β° FoV, 640Γ480 @ 120 FPS, fixed exposure.
- Illumination: 8Γ 4000K white LEDs (2835 package) in an annular PCB.
- Elastomer: PDMS base + Ecoflex 00-10 layer with 33 visual markers for shear-force detection.
- Tactile vector field: V β β^{vΓ4} where v = 33 (markers); channels are (normal force magnitude, shear direction-x, shear direction-y, shear magnitude).
- Marker tracking via Farneback optical flow (FarnebΓ€ck 2003); depth reconstruction via grayscale β indentation-depth LUT calibrated with a 2 mm spherical indenter.
- Data collected: 20 objects, ~300k frames @ 30 FPS.
The exoskeleton is isomorphic to the robot hand, deliberately minimising the domain gap.
4.1 Establish-Contact (Flow Matching).
- Input: object point cloud D, target pose (P^obj_target, Ο^obj_target).
- Output: (A^hand_0, R^wrist_0, T^wrist_0).
- Backbone: PointNet++ for point-cloud features, FiLM conditioning into a 5-layer MLP (widths 128, 128, 512, 1024, 1024), flow-matching loss
L_contact = E[β(X_1 β X_0) β u(X_t, t, cond)βΒ²]with X_t = (1βt)X_0 + tX_1. - Training: batch 128, AdamW, LR 1e-4, 1.3M steps.
- FM chosen over diffusion policy for faster training and inference.
4.2 DexMove-Policy (transformer flow matching).
- State at time t (past Tp = 5 frames):
(P^hand, A^hand, R^wrist, T^wrist, P^obj, Ο^obj, C, G)_{βTp:0}where P^hand β β^{JΓ3} joint positions, C β β^{FΓ3} per-finger contact positions in local frame, G β β^F per-finger pressing force. - Target: future hand state X_1 over Tf = 5 frames.
- Architecture: Past tokens + global target token (linear projection of (P_target, ΞΈ_target)) + continuous time token (Fourier features β MLP) β Transformer encoder produces memory M β β^{(Tp+2)Γd}. The noised state X_t and planned future force G_{1:Tf} are FiLM-fused into query tokens for the Transformer decoder, which outputs the velocity field.
- Training: 200k steps, batch 2048; then 20k steps of ReFlow (Liu et al. 2022) to compress to a 10-step inference sampler.
- Inference: ~22 ms/chunk on RTX 4090, executed at 30 Hz.
4.3 TaFo-Net (tactile force planner). Given target pose, Tp = 5 past frames of object states, and per-finger tactile vector fields V_{βTp:0} β β^{TpΓFΓvΓC}, predict future tactile fields V_{1:Tf} from which per-finger forces G_{1:Tf} are extracted. Three-stage architecture:
- Per-finger spatial encoding. Each V_{t,f} β β^{vΓC} β token U_{t,f} via a lightweight transformer with learnable + geometry-informed marker positional embeddings.
-
Cross-finger attention. For each frame i, multi-head self-attention across F fingers augmented with per-finger type embeddings g_f:
Ε¨_{i,1:F} = CF(U_{i,1:F} + g_{1:F}). - Finger-wise causal temporal attention. Causal mask so a query at i can only attend to tokens at times β€ i, preventing future leakage.
Training loss L_rec = Ξ£_t Ξ£_f βVΜ_{t,f} β V_{t,f}βΒ², with random dropout of time steps, fingers, and markers for robustness.
- Robot: Franka FR3 + Allegro Hand. Position control through ROS β Cartesian for the arm, joint-space PID for the hand.
- Vision: 3Γ Realsense D435i depth cameras (one near elbow, two on opposite sides). Object pose via ArUco markers (calibrated; markerless results also reported using FoundationPose).
- Compute: RTX 4090 deployment.
Six everyday objects: LEGO, mouse, keyboard, book, large can, small can. Two surfaces: Friction A (clean table) and Friction B (tape strips, unseen during data collection). Initial-yaw-error bins:
| Method | 0Β° < Ο_target < 30Β° (A / B) | 30Β°β60Β° (A / B) | 60Β°β90Β° (A / B) |
|---|---|---|---|
| Open-loop | 36.7 / 10.0 | 23.3 / 0.0 | 3.3 / 0.0 |
| DyWA (Lyu et al. 2025) | 50.0 / 36.7 | 46.7 / 30.0 | 50.0 / 33.3 |
| CORN (Cho et al. 2024) | 43.3 / 36.7 | 46.7 / 40.0 | 43.3 / 43.3 |
| DexMove | 86.7 / 86.7 | 80.0 / 83.3 | 70.0 / 60.0 |
DexMove maintains performance under unseen friction (gap is small, A vs. B); DyWA and CORN show pronounced degradation, reflecting their sensitivity to spatial friction variability.
| Method | 0β15 cm | 15β30 cm | 30β45 cm |
|---|---|---|---|
| DyWA | 36.1 | 52.2 | 60.6 |
| CORN | 41.4 | 54.5 | 62.1 |
| DexMove | 8.3 | 10.9 | 12.4 |
DexMove is 3-5Γ faster because multi-finger contact reduces the number of action primitives needed to reach a target pose.
Across the six objects, average success 77.8%; +36.6% over ablated baselines, ~300% efficiency improvement over baseline methods (i.e., ~4Γ speed).
| Object | Trials | Success |
|---|---|---|
| Rag doll | 30 | 96.7% |
| Tissue packet | 30 | 100% |
Deformability actually helps β compliant contact stabilises contact formation.
Random stacking of objects beneath the manipulated items. Tested on book, large can, LEGO with two conditions (w/o finetune, w/ finetune on 15 min of uneven-surface tactile data + masked contact intervals). Specific numbers not stated in extracted body but the paper claims robustness with light finetuning.
Replacing ArUco markers with FoundationPose: success rates of 16.7, 13.3, 93.3, 96.7, 60.0, 76.6% for the six objects. Hand-occlusion-induced pose estimation errors hurt small objects the most.
| Noise Ο | Err-MSE (book) | SR (book) | Err-MSE (can) | SR (can) |
|---|---|---|---|---|
| 0 | 0.0112 | 90.0% | 0.0351 | 63.3% |
| 0.05 | 0.0108 | 86.7% | 0.0615 | 56.7% |
| 0.1 | 0.0415 | 80.0% | 0.1239 | 43.3% |
| 0.2 | 0.1721 | 53.3% | 0.3005 | 20.0% |
| 0.4 | 0.3219 | 13.3% | 0.5312 | 3.3% |
Robust up to Ο = 0.1 (sensor noise + TaFo-Net prediction errors). Attributed to (i) noisy training tactile signals, (ii) random dropout during training.
The learned policy generalizes to language-conditioned, long-horizon tasks β one of the paper's three stated significance pillars. Three demonstrated scenarios: (i) structured sorting ("move box A to region 1"); (ii) language-driven humanβmachine collaboration, where a VLM (SoFar, Qi et al. 2025) converts a natural-language command (e.g., "put the grip of the electric drill into a person's hand") into a 3-DoF target pose fed to the policy for a non-prehensile handover; (iii) desktop tidying, relocating each item to an assigned position from a predefined layout.
| Method | LEGO | Mouse | Book | Keyboard | Large Can | Small Can |
|---|---|---|---|---|---|---|
| Wrist-Only (auto, fingers locked) | 13.3 | 0.0 | 33.3 | 20.0 | 0.0 | 0.0 |
| Wrist-Only* (teleoperated) | 0.0 | 73.3 | 100.0 | 100.0 | 6.7 | 10.0 |
| w/o Cross-Finger | 13.3 | 3.3 | 63.3 | 50.0 | 0.0 | 3.3 |
| w/o Shear-Force | 70.0 | 66.7 | 33.3 | 13.3 | 0.0 | 0.0 |
| w Heuristic Force | 36.7 | 43.3 | 66.7 | 0.0 | 0.0 | 0.0 |
| DexMove (Ours) | 66.7 | 86.7 | 90.0 | 90.0 | 63.3 | 70.0 |
Findings:
- Wrist-only can handle flat/planar objects (book, keyboard) when teleoperated β confirming the wrist contributes substantially β but rarely succeeds on objects requiring non-coplanar fingertip contacts (LEGO, mouse) or shape-induced grasp adjustments (cans).
- Cross-finger attention is essential β without it TaFo-Net fails on round/heavy objects (large can goes from 63.3 β 0%).
- Shear force is critical for slip detection and heavier objects. Without it, the model converges on smoothed averaged states β fine for light objects (LEGO, mouse) but fails on cans.
- Heuristic force (slip-detection + force increment, Γ la Lin et al. 2025) cannot replace TaFo-Net β incrementing force after slip is reactive, not predictive.
The authors explicitly list:
- Articulated objects. Objects with movable parts (e.g., a telephone with a handset) shift during manipulation and destabilise contact.
- Spherical objects. Tend to roll; stable initial contact is difficult; slippage risk increases.
- Hand-pose failure modes. Some grasps cause the object to topple β e.g., pushing a tall can while gripping only its lid causes a tip-over and restricts rotational motion.
Additional implicit limitations from experiments:
- Markerless pose estimation degrades small-object performance (16.7-13.3% on LEGO and mouse vs. 90% with ArUco).
- The simulation-to-real bridge depends on the isomorphic exoskeleton; non-isomorphic robot hands would require new data collection.
- Friction generalisation tested in only two regimes (clean vs. tape).
- Authors plan to extend to prehensile + non-prehensile integration in future work.
DexMove opens the non-prehensile regime for multi-fingered hands by:
- Treating wrist and fingers as a coupled control problem rather than two separate stages β the wrist-motion regulariser in IK biases trajectories toward finger-driven manipulation.
- Bringing real tactile data into the loop via the isomorphic wearable with R-Tac sensors, sidestepping the layout mismatch of tactile gloves.
- Using flow matching (faster than diffusion policy) for both the contact policy and the trajectory policy, with ReFlow compression to 10 inference steps.
The skill also plugs into language-conditioned, long-horizon pipelines (object sorting, VLM-driven handover via SoFar, desktop tidying β Β§5.5), showing the non-prehensile primitive composes into higher-level task planning.
Versus related work:
- CORN (Cho et al. 2024) & DyWA (Lyu et al. 2025): gripper-based non-prehensile baselines that lose accuracy on cylindrical objects (single contact point). DexMove leverages distributed multi-finger contacts.
- DemoGrasp / DexNDM: grasp-centric dexterous learning. Complementary β DexMove handles the non-grasp regime they exclude.
- DexUMI: wearable dexterous teleoperation for grasping. DexMove extends the isomorphic-wearable idea to force-aware non-prehensile data.
- 3D-ViTac (Huang et al. 2024), Higuera et al. 2025: tactile-rich manipulation. DexMove adds the multi-contact non-prehensile angle and the wrist-finger synergy claim.
- DexNoMa (Li et al. 2025b): dexterous non-prehensile manipulation, but on smaller object sets and without tactile sensing.
The 77.8% real-world average across six varied objects, with 3-5Γ faster execution than gripper baselines and 96-100% on deformable objects, is the strongest quantitative argument for why multi-fingered hands matter for non-prehensile tasks β they are not just more dexterous, they are more stable and more efficient than the gripper alternatives.
- OpenReview: https://openreview.net/forum?id=dT3ZciXvNX
- Project: https://peilin-666.github.io/projects/DexMove/
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)