-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 InstructVLA
Venue: ICLR 2026 Β· arXiv: 2507.17520 Β· GitHub: InternRobotics/InstructVLA Affiliations: Shanghai AI Lab, USTC, Zhejiang University Β· Base VLM: Eagle2-2B Category: Training Approach Trend tag: Trend 2
flowchart TB
Stage1[Stage 1: Action pretraining<br/>heterogeneous manipulation data<br/>actions + rule-based language motion] --> Stage2[Stage 2: VLA-IT<br/>freeze action expert,<br/>add language LoRA + scale head]
Stage2 --> MoE{MoE-adaptation<br/>LoRA experts in LLM}
MoE -- scale head predicts gating Ξ»α΅’ --> Blend[Adaptively blend expert outputs<br/>reasoning β action]
Data[VLA-IT 650K dataset<br/>instructions, captions, QA pairs] --> Stage2
Like Actions as Language: existing VLA models sacrifice multimodal reasoning for task-specific manipulation and suffer catastrophic forgetting of pretrained vision-language capability. InstructVLA aims to preserve the flexible reasoning of the base VLM while delivering leading manipulation performance, using embodied reasoning to help action.
- Base VLM: Eagle2-2B backbone.
- Two-stage pipeline (order matters): Stage 1 = action pretraining on heterogeneous manipulation data, jointly predicting actions (flow-matching objective) and rule-based annotated language motion (LM loss). Stage 2 = Vision-Language-Action Instruction Tuning (VLA-IT), which freezes the action expert and adds a new language LoRA adapter plus a scale head of the MoE-adaptation.
- MoE-adaptation (not a dual system): LoRA modules act as experts inside the LLM backbone. A scale head predicts gating coefficients Ξ»α΅’ per expert by classifying the hidden state, adaptively blending their outputs to switch between reasoning and action generation. It does not route separate token types to separate experts; a single VLM emits both text and latent actions.
- VLA-IT 650K: 650K human-robot interactions annotated with diverse instructions, scene captions, and grounded QA pairs, trained jointly with standard VLM corpora.
- In-domain SimplerEnv: +33% over SpatialVLA.
- SimplerEnv-Instruct (newly introduced 80-task benchmark requiring closed-loop control + high-level instruction understanding): outperforms a fine-tuned OpenVLA by 96% and an action expert aided by GPT-4o by 29%.
- Surpasses baseline VLMs on multimodal tasks and shows inference-time scaling β textual reasoning boosts manipulation in sim and real.
Contributes both a method (MoE-adaptation that preserves VLM reasoning during action learning) and a dataset (VLA-IT 650K) reusable by other groups. Demonstrates that embodied reasoning can be co-trained with action generation without sacrificing pretrained multimodal capability, enabling steerable instruction following.
- Actions as Language (the data-relabeling counterpart)
- Embodied-R1 (RL-based reasoning)
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)