Skip to content

IROS 2026 AtomVLA

hwoo.han edited this page Sep 7, 2026 · 2 revisions

IROS 2026 β€” AtomVLA: Scalable Post-Training for Manipulation via Predictive Latent World Models

Venue: IROS 2026 (Pittsburgh) Β· paper #33 Β· Huazhong Univ. of S&T Β· HKU Β· Tsinghua Β· Beihang (Sun, Xu, Cao, … Chen). Paper: arXiv 2603.08519 (Mar 2026) Β· datasets/checkpoints/code to be released. The WAMΓ—VLA datapoint of IROS 2026 β€” a latent world model used to score a VLA's action chunks during offline post-training, closing the instruction-grounding gap on long-horizon tasks. Companions: World Models Β· VLA Hybrid Architectures Β· Multi-Task VLA Β· DYNA-2 Β· IROS 2026 survey.

AtomVLA's two-stage pipeline β€” Stage I SFT (Qwen3-VL backbone + flow-matching action head); Stage II post-training: the frozen VLA rolls out k candidate action chunks in the environment, a World-Model encoder scores them against subtask/goal frames, and the resulting reward drives offline GRPO on the action head (architecture figure from Sun et al., arXiv 2603.08519, Β© the authors)

1. Problem

VLAs execute complex multi-step behaviors better with robust instruction grounding, but current SFT relies on coarse high-level instructions β€” no explicit intermediate guidance β€” so long-horizon tasks suffer compounding errors. The paper frames this as an instruction-grounding gap and seeks a scalable post-training recipe to close it.

2. Method

AtomVLA is "the first subtask-aware VLA framework integrated with a scalable offline post-training pipeline."

  • Atomic-subtask decomposition β€” an LLM decomposes high-level demonstrations into fine-grained atomic subtasks (the intermediate guidance the SFT signal lacks).
  • Latent world-model scoring β€” a pretrained predictive world model scores candidate action chunks against subtask goals in latent space, mitigating error accumulation and improving long-horizon robustness. (This is the reconstruction-free, deployable latent-WAM flavor β€” cf. Ο‰-0/DYNA-2 β€” used as an evaluator/critic, not an in-path generator.)
  • Offline GRPO β€” the latent scoring enables Group Relative Policy Optimization without online physical-robot rollouts (the expensive part of RL-for-VLA β€” cf. DyGRO-VLA, RL for VLA).

3. Results

  • LIBERO 97.0%, LIBERO-PRO 48.0% avg success vs baselines; strong robustness under perturbations.
  • Real-world on the Galaxea R1 Lite β€” broad applicability, especially long-horizon tasks.
  • Datasets / checkpoints / code to be publicly released.

4. Why it matters (WAM lens)

AtomVLA is a clean instance of IROS 2026's WAM trend (survey Β§5.1): the world model has moved from pixel generator to latent evaluator inside a VLA post-training loop. It combines three threads the wiki tracks β€” latent WAM (World Models), subtask/instruction grounding (Multi-Task VLA, AtomVLA's "instruction gap" β‰ˆ LangForce/DISC's language-shortcut problem), and offline RL for VLA (DyGRO-VLA). The offline-GRPO-via-latent-scoring move is the notable systems contribution: RL-quality post-training without robot rollouts.

Limitations (reviewer): results are LIBERO/LIBERO-PRO sim + one real platform; the world model's scoring fidelity is the ceiling (a bad latent critic misranks chunks); LLM subtask decomposition quality is an unmeasured dependency.

5. Links

← Back to IROS 2026 survey Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally