Skip to content

ICLR 2026 InstructVLA

Heungwoo edited this page Jun 1, 2026 · 2 revisions

InstructVLA β€” Vision-Language-Action Instruction Tuning

Venue: ICLR 2026 Β· arXiv: 2507.17520 Β· GitHub: InternRobotics/InstructVLA Affiliations: Shanghai AI Lab, USTC, Zhejiang University Β· Base VLM: Eagle2-2B Category: Training Approach Trend tag: Trend 2

Approach diagram

flowchart TB
  Stage1[Stage 1: Action pretraining<br/>heterogeneous manipulation data<br/>actions + rule-based language motion] --> Stage2[Stage 2: VLA-IT<br/>freeze action expert,<br/>add language LoRA + scale head]
  Stage2 --> MoE{MoE-adaptation<br/>LoRA experts in LLM}
  MoE -- scale head predicts gating Ξ»α΅’ --> Blend[Adaptively blend expert outputs<br/>reasoning ↔ action]
  Data[VLA-IT 650K dataset<br/>instructions, captions, QA pairs] --> Stage2
Loading

Problem

Like Actions as Language: existing VLA models sacrifice multimodal reasoning for task-specific manipulation and suffer catastrophic forgetting of pretrained vision-language capability. InstructVLA aims to preserve the flexible reasoning of the base VLM while delivering leading manipulation performance, using embodied reasoning to help action.

Method

  • Base VLM: Eagle2-2B backbone.
  • Two-stage pipeline (order matters): Stage 1 = action pretraining on heterogeneous manipulation data, jointly predicting actions (flow-matching objective) and rule-based annotated language motion (LM loss). Stage 2 = Vision-Language-Action Instruction Tuning (VLA-IT), which freezes the action expert and adds a new language LoRA adapter plus a scale head of the MoE-adaptation.
  • MoE-adaptation (not a dual system): LoRA modules act as experts inside the LLM backbone. A scale head predicts gating coefficients Ξ»α΅’ per expert by classifying the hidden state, adaptively blending their outputs to switch between reasoning and action generation. It does not route separate token types to separate experts; a single VLM emits both text and latent actions.
  • VLA-IT 650K: 650K human-robot interactions annotated with diverse instructions, scene captions, and grounded QA pairs, trained jointly with standard VLM corpora.

Results

  • In-domain SimplerEnv: +33% over SpatialVLA.
  • SimplerEnv-Instruct (newly introduced 80-task benchmark requiring closed-loop control + high-level instruction understanding): outperforms a fine-tuned OpenVLA by 96% and an action expert aided by GPT-4o by 29%.
  • Surpasses baseline VLMs on multimodal tasks and shows inference-time scaling β€” textual reasoning boosts manipulation in sim and real.

Significance

Contributes both a method (MoE-adaptation that preserves VLM reasoning during action learning) and a dataset (VLA-IT 650K) reusable by other groups. Demonstrates that embodied reasoning can be co-trained with action generation without sacrificing pretrained multimodal capability, enabling steerable instruction following.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally