-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 GuidedVLA
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: VLA Models Β· paper #84 Authors: Xiaosong Jia, Bowen Yang, Zuhao Ge, Xian Nie, Yuchen Zhou, Cunxin Fan, Yufeng Li, Yilin Chai, Chao Jing, Zijian Liang, Qingwen Bu, Haidong Cao, Chao Wu, Qifeng Li, Zhenjie Yang, Chenhe Zhang, Hongyang Li, Zuxuan Wu, Junchi Yan, Yu-Gang Jiang ...more> arXiv: 2605.12369 Β· program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Top: benchmark gains over the Ο0 baseline on LIBERO-Plus, RoboTwin 2.0 (77.38β90.63) and two real-world platforms. Bottom panels contrast baseline vs. GuidedVLA on the three specialized factors: skill recognition (correct MoveβSweepβDump sequencing), object grounding (attention concentrated on the target object instead of scattered), and geometry perception (clean depth prediction).
Existing VLAs learn task-relevant features only implicitly through end-to-end action supervision, so the action decoder often latches onto spurious correlations β visual shortcuts, background noise β that hurt out-of-domain generalization. GuidedVLA (Fudan TEAI, SJTU, OpenDriveLab/HKU) asks whether explicitly guiding the action decoder toward task-relevant factors fixes this.
The core idea treats the action decoder not as a monolithic learner but as an assembly of functional components: individual cross-attention heads are supervised with auxiliary signals while the remaining heads stay free. The initial instantiation (on a Ο0 base policy) uses three specialized heads: an object head whose attention is constrained to Grounded-SAM-annotated target/destination regions; a skill head aligning internal features with temporal sub-skill phases; and a depth head distilling features from a depth encoder (no depth annotation needed β injected architecturally). A largely automatic annotation pipeline (Qwen3-VL + SAM2, ~11Γ faster than manual: ~4 min vs. ~43.5 min per 50 episodes; 95.2% auto accuracy for object masks, 87.3% for skill labels) produces the guidance data.
On LIBERO-Plus, GuidedVLA reaches 75.4% average vs. 68.2% for Ο0, topping 12 baselines including OpenVLA-OFT (69.6%) and DreamVLA (69.9%); single-head ablations show the object head strongest on the Object suite (82.5%, +8.4 over Ο0), skill head best on Goal (68.9%), depth head best on Spatial. On RoboTwin 2.0 the full model lifts average success from 77.38% to 90.63% (e.g., Click Bell 35%β63% with the depth head, Beat Hammer Block 78%β96%). On real robots (ALOHA AgileX household tasks; PSI-Bot RealMan chemistry-lab tasks, 20 trials each) it beats the base policy in every setting: 75.8% vs. 55.8% in-domain, 67.5% vs. 44.2% with scene distractors, 79.2% vs. 57.5% under lighting shifts. Factor-quality analyses show success rises monotonically with head quality (e.g., depth-feature ratio: 15.0%β74.2%).
A plug-and-play middle path between fully implicit end-to-end VLAs and hand-built modular stacks: supervise a few attention heads, keep the rest free. The demonstrated correlation between factor quality and success supports interpretable-by-construction action decoding β relevant to Review-VLA-Architecture and the robustness themes of Review-VLA-Evaluation.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)