Skip to content

NeurIPS 2025 ChatVLA 2

Heungwoo edited this page Jun 1, 2026 · 2 revisions

ChatVLA-2 β€” VLA with Open-World Embodied Reasoning

Venue: NeurIPS 2025 Β· Authors: Midea Group + East China Normal University Β· arXiv: 2505.21906 Category: VLA Architecture / CoT Reasoning

Approach diagram

flowchart LR
  V[Vision] --> VLM[VLM backbone]
  L[Language] --> VLM
  VLM --> MoE[Dynamic MoE:<br/>reasoning experts βŠ• action experts]
  MoE -- reasoning path --> R[VLM-preserved reasoning output]
  MoE -- action path --> AE[Action decoder]
  AE --> A[Actions]
  note[2-stage pipeline:<br/>1. co-train image-text βŠ• robot data<br/>2. freeze VLM, train action expert only]
Loading

Problem

Fine-tuning a VLM into a VLA nearly always destroys open-world reasoning β€” the model gets good at its training tasks but loses the broad commonsense / math / language-following that its VLM pretraining provided. VLM4VLA documents the phenomenon; ChatVLA-2 attacks the root cause.

Built on DexVLA as the foundational architecture, with Qwen2-VL as the core VLM. Because Qwen2-VL's LLM lacks native MoE support, the authors insert a dynamic mixture-of-experts inside the VLM backbone to disentangle multimodal-understanding vs. robotic-action feature spaces without disrupting the pretrained structure:

  • 8 experts total, top-2 selected per inference via an adaptive gating network conditioned on the visual/textual input;
  • Some experts capture shared multimodal/spatial-reasoning features, others specialize on task-specific (manipulation) features;
  • A separate action expert (inherited from DexVLA) decodes actions, aligned with the model's internal reasoning.

Two-stage training pipeline (paper Β§3.3):

  1. Open-world reasoning stage β€” co-train on image-text data βŠ• robot data, simultaneously training robotic actions, to preserve pretrained multimodal knowledge and establish reasoning↔action connections.
  2. Reasoning-following stage β€” freeze the entire VLM and train only the action expert, so open-world reasoning is preserved while instruction/reasoning-following in action execution is enhanced.
  • In-domain manipulation: ChatVLA-2 performs comparably to strong imitation baselines (11/13 vs DexVLA 12/13 and Ο€0 12/13 on the math-matching task; all near-saturated). The paper is explicit that it "does not significantly outperform models like Ο€0 and DexVLA" in-domain.
  • Open-world is where it wins decisively. On the math-matching game (robot reads whiteboard equations and picks matching number cards) ChatVLA-2 reaches an 82.7% manipulation success rate (43/52) with OCR score 3.58 and math-reasoning score 1.73, while baselines (Octo, Diffusion Policy, OpenVLA, DexVLA, Ο€0) score 0–10/52 β€” these reasoning/OCR abilities were never explicitly trained in the VLA. A separate toy-placement task shows analogous open-world spatial-reasoning gains over unseen objects/directions.
  • Ablation (paper Table 4): removing Stage 2 collapses open-world control to ~23%; removing Stage 1 drives open-world reasoning to near-zero β€” both stages are necessary (reasoning is generated in Stage 1 but only injected into action execution in Stage 2).

Significance

One of the strongest published demonstrations that MoE is an effective mechanism for the reasoning-vs-action feature conflict β€” and that simply freezing the VLM in a second stage suffices to preserve open-world reasoning while injecting it into control. Seeds the ICLR 2026 MoE cluster (HiMoE-VLA, AdaMoE). Note it is a single-system, MoE-in-backbone design (not an explicit fast/slow dual-system like Fast-in-Slow or ThinkAct), though it is frequently grouped with them as a reasoning-preserving VLA.

Links

Related pages

← Back to NeurIPS-2025

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally