Skip to content

RSS 2026 TactAlign

hwoo.han edited this page Aug 9, 2026 · 2 revisions

TactAlign: Human-to-Robot Policy Transfer via Tactile Alignment

Venue: RSS 2026 (Sydney, Jul 13–17) Β· Session: Manipulation 1 Β· paper #6 Authors: Youngsun Wi, Jessica Yin, Elvis Xiang, Akash Sharma, Jitendra Malik, Mustafa Mukadam, Nima Fazeli, Tess Hellebrekers arXiv: 2602.13579 Β· program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

TactAlign pipeline (Figure 1 of arXiv 2602.13579, Β© the authors)

Left: unpaired human (tactile-glove) and robot demonstrations feed two stages β€” tactile self-supervised learning, then cross-sensor tactile alignment via rectified flow that maps glove latents into the robot tactile latent space. Right: the resulting human-to-robot policies shown on real hardware β€” object-level generalization with <5 min of human demos, generalization beyond demonstrated object instances, dexterous manipulation from human data only, and task-level generalization.

Problem

Human demonstrations collected with wearable tactile gloves are fast and naturally dexterous, but transferring the tactile signals to a robot is hard because sensors and embodiments differ. Existing human-to-robot (H2R) approaches that use touch typically assume identical tactile sensors, need paired data, or tolerate little embodiment gap; the concurrent UniTacHand instead requires strict spatiotemporal human–robot pairing, which is impractical during sliding contact or dynamic object motion.

Method

TactAlign aligns human and robot tactile observations in a shared latent space from unpaired demonstrations of the same task. Stage 1: separate human and robot tactile encoders are pretrained with self-supervised learning (JEPA-style architecture with MSE reconstruction and cross-attention pooling to a fixed-dimensional latent), handling very different resolutions (OSMO glove 1Γ—3 vs. Xela robot skin 30Γ—3). Stage 2: a rectified flow learns a velocity field transporting glove latents to robot latents, trained on noisy pseudo-pairs mined from hand–object interaction similarity (matching poses and pose deltas of transitions under a threshold Ξ΄), with no paired datasets, manual labels, or privileged information; sampling uses plain Euler integration. Hardware: OSMO tactile glove on the human side; a Franka Emika Panda with Xela sensing and a RealSense D455 on the robot side.

Results

On H2R co-training across pivoting, insertion, and lid closing, TactAlign reaches 76%, 72%, and 74% success respectively β€” 100% on seen-by-both objects, 71% on human-only objects, and 65.5% on unseen-by-both objects (averaged over the three tasks). Ablations: removing tactile input costs βˆ’59% average success (up to βˆ’100% on pivoting); removing alignment (raw tactile features) costs βˆ’51% on average and causes near-complete failure on seen objects for pivoting/insertion. On the zero-shot dexterous light-bulb-screwing task trained from human data only, TactAlign achieves 100% success (~61 s to illumination) while both the no-tactile and no-alignment baselines score 0%.

Significance

A rare demonstration that tactile signals β€” not just vision or kinematics β€” can be transferred across the human–robot embodiment gap without paired data, using flow matching as a cross-sensor bridge. Connects directly to the wiki threads Review-Tactile-VLA and Review-Dexterous-Manipulation.

← Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally