Skip to content

RSS 2026 SID

hwoo.han edited this page Aug 9, 2026 · 2 revisions

SID: Sliding into Distribution for Robust Few-Demonstration Manipulation

Venue: RSS 2026 (Sydney, Jul 13–17) Β· Session: Manipulation 1 Β· paper #8 Authors: Yicheng Ma, Wei Yu, Zhian Su, Xidan Zhang, Huixu Dong arXiv: 2605.13428 Β· program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

SID concept: motion field slides OOD states back into distribution (Figure 1 of arXiv 2605.13428, Β© the authors)

The teaser shows SID's two-regime view of a manipulation episode: gripper poses starting out-of-distribution follow a learned object-centric motion field (black arrows) that "slides" them into the demonstrated in-distribution region (dashed funnel), where a lightweight egocentric execution policy (blue arrows) takes over for the contact-rich interaction β€” illustrated by the energy-landscape inset whose gradient vanishes near the demonstration manifold.

Problem

Few-demonstration visuomotor policies fail mostly through distribution shift: with low-coverage data, test-time object poses, viewpoints, and disturbances put the robot in states the demonstrations never covered. The authors observe that an episode has two qualitatively different regimes β€” an underdetermined approach phase and a locally sensitive, interaction-rich execution phase β€” and that a single end-to-end policy must resolve both at once, making it brittle under pose shifts and perturbations.

Method

SID (Grasp Lab, Zhejiang University + Torch Kernel Co.) factorizes control into four components: (i) an object-centric motion field f_ΞΈ learned from canonicalized approach-phase demonstrations, implemented as a gradient-descent-style dynamical system over a pose-aligned SE(3) potential that produces large corrective motions far from the demonstration manifold and vanishes near convergence; (ii) a kinematically consistent point-cloud reprojection augmentation that perturbs the end-effector pose, reprojects segmented wrist-camera point clouds via fixed hand–eye calibration, and updates relative actions to preserve action–observation consistency (also generating ID/OOD labels); (iii) an egocentric execution policy trained with conditioned flow matching on wrist-centric point clouds and gripper width; and (iv) two inference pipelines β€” open-loop handoff (SID-O) and a closed-loop variant (SID-C) whose auxiliary ID-confidence head triggers field-based re-alignment when observations drift OOD. Canonicalization uses 6D pose estimation and SAM2/SAM3 object segmentation. Each task uses only two raw demonstrations, expanded to 100 training samples by augmentation.

Results

On six real-world tasks (Open Drawer, Pour Water, Hang Tape, Hang Cup, PnP-Box, Multi-PnP-Box; 50 trials each), SID-C reaches 86–92% success under OOD initializations with 2 demos, versus Ο€0.5 at 0–14% OOD (trained with 100 demos) and retrieval baselines MT3/Ret-BC at 36–78% (10 demos); ACT and DP3 largely collapse OOD. In the dynamic-disturbance setting SID-C scores 82–88% across four tasks, beating the strongest baseline Ο€0.5 (68–80%), and both SID variants stay reliable in cluttered scenes and on recomposed long-horizon tasks built by reusing learned sub-skills. The abstract's headline: ~90% OOD success with two demonstrations and under a 10% drop with distractors and external disturbances.

Significance

A clean articulation of "online distribution recovery" as an alternative to scaling data: instead of covering the state space, steer the system back to where the policy is competent. Complements the equivariance and coarse-to-fine lines discussed in Review-Dexterous-Manipulation, and its OOD-confidence-gated re-alignment echoes the runtime-monitoring themes in Review-VLA-Evaluation.

← Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally