Skip to content

RSS 2026 LAP

hwoo.han edited this page Aug 10, 2026 · 2 revisions

LAP: Language-Action Pre-training Enables Zero-Shot Cross-Embodiment Transfer

Venue: RSS 2026 (Sydney, Jul 13–17) Β· Session: Imitation learning 3 Β· paper #203 Authors: Lihan Zha, Asher James Hancock, Mingtong Zhang, Tenny Yin, Yixuan Huang, Dhruv Shah, Allen Z. Ren, Anirudha Majumdar (Princeton University Β· Physical Intelligence) arXiv: 2602.10556 Β· program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Language-Action Pre-training overview and zero-shot generalization panel (Figure 1 of arXiv 2602.10556, Β© the authors)

Figure 1: Overview of LAP. Low-level end-effector actions (e.g. "Move forward 3cm", "tilt forward 30 degrees") are expressed directly as natural-language "language-actions" and used to supervise a PaliGemma-3B/Gemma3 backbone, unifying robot action prediction with motion-prediction VQA in a "Semantically Unified Action Space." The right panels show zero-shot generalization to unseen embodiments (SOTA VLA 0% vs. LAP-3B 52%) and transferable-representation scaling curves.

Problem

Despite large-scale multi-embodiment pre-training, existing VLAs stay tightly coupled to their training embodiments and rarely work zero-shot on new robots β€” even minor differences (an altered gripper or wrist-camera location) usually demand costly per-embodiment fine-tuning.

Method

Language-Action Pre-training (LAP) represents low-level robot actions directly in natural language ("language-actions"), aligning action supervision with the pre-trained VLM's input–output distribution. It requires no learned tokenizer, no costly annotation, and no embodiment-specific architecture; language-actions are parsed from raw actions under a fixed coordinate convention. The instantiation, LAP-3B, initializes its VLM backbone from PaliGemma-3B and uses a Mixture-of-Transformers architecture combining the LAP-trained VLM with a lightweight diffusion-based action expert for real-time control (following Ο€0.5, from which it differs only in action representation). It is trained on open-sourced robot datasets including Open X-Embodiment, and can additionally co-train with VQA data.

Results

Across multiple novel robot embodiments and manipulation tasks, LAP-3B attains over 50% average zero-shot success (52% reported) β€” roughly a 2Γ— improvement (~30 points absolute) over the strongest prior VLAs β€” while all open-sourced VLA baselines collapse to near-zero zero-shot success. It consistently outperforms replicated Ο€0.5 and Ο€0 baselines by about 15 percentage points, and the paper reports efficient adaptation, favorable scaling, and additional gains from unifying action prediction with VQA through co-training.

Significance

Presented as the first VLA to achieve substantial zero-shot transfer to unseen real embodiments without embodiment-specific fine-tuning, advancing cross-embodiment generalist policies discussed in Review-Human-Video-Transfer and Review-Realtime-Execution.

← Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally