Repository navigation
RSS 2026 ViTacFormer
Venue: RSS 2026 (Manipulation session) Β· Authors: Liang Heng, Haoran Geng, Kaifeng Zhang, Pieter Abbeel, Jitendra Malik (Berkeley) Β· arXiv: 2506.15953 Β· project Category: Visuo-tactile representation learning for dexterous hands Trend tag: RSS 2026 thread 4 β contact as representation Hardware: dual Realman arms + two SharpaWave anthropomorphic hands (5 digits, 17 DoF each) with high-resolution 320Γ240 tactile sensors on all 10 fingertips; exoskeleton-glove + VR teleoperation
Compiled from the verified RSS 2026 abstract and the paper's Β§IIIβIV / Fig. 2.

Figure 2 of the paper β the model is a conditional variational auto-encoder. Left: a transformer encoder maps the action sequence + proprioception to a style variable z (CVAE prior, ACT-style). Right: the ViTacFormer decoder consumes z, joints, multi-camera images, and touch through cross-attention, and β the key design β carries a future-touch-prediction pathway: the predicted tactile signal is fed back as an input alongside real touch while the model auto-regressively generates the action sequence. Contact anticipation is thus built into the representation the policy acts from, not appended as an observation.
Vision-based dexterous manipulation fails exactly where it matters β fine-grained control under occlusion and contact. Tactile signals are usually appended as extra observations rather than fused into a representation that anticipates contact.
- Cross-attention encoder fusing high-resolution vision and touch into a shared latent.
- Autoregressive tactile-prediction head: the representation is trained to forecast future contact signals, not just encode current ones β making anticipated contact part of the state.
- Easy-to-challenging curriculum progressively refines the visuo-tactile latent space.
- The learned representation drives imitation learning on multi-fingered anthropomorphic hands.
- β50% higher success rates than prior state-of-the-art across challenging real-world benchmarks.
- First system (per the authors) to autonomously complete long-horizon dexterous tasks of up to 11 sequential stages, sustaining 2.5 minutes of continuous operation with an anthropomorphic hand.
The flagship of RSS 2026's contact-as-representation thread: where Contact-Grounded Policy grounds contact through prediction-to-control mapping, ViTacFormer grounds it through predictive representation β the tactile analogue of world-model forecasting. The 11-stage/2.5-minute result sets the long-horizon bar for anthropomorphic-hand autonomy. Slots into Review-Tactile-VLA's "predict-touch" family and Review-Dexterous-Manipulation's IL branch; the Berkeley (Abbeel/Malik) lineage connects it to the broader hand-scaling agenda.
β RSS 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)