-
Notifications
You must be signed in to change notification settings - Fork 0
IROS 2026 AtomVLA
Venue: IROS 2026 (Pittsburgh) Β· paper #33 Β· Huazhong Univ. of S&T Β· HKU Β· Tsinghua Β· Beihang (Sun, Xu, Cao, β¦ Chen). Paper: arXiv 2603.08519 (Mar 2026) Β· datasets/checkpoints/code to be released. The WAMΓVLA datapoint of IROS 2026 β a latent world model used to score a VLA's action chunks during offline post-training, closing the instruction-grounding gap on long-horizon tasks. Companions: World Models Β· VLA Hybrid Architectures Β· Multi-Task VLA Β· DYNA-2 Β· IROS 2026 survey.

VLAs execute complex multi-step behaviors better with robust instruction grounding, but current SFT relies on coarse high-level instructions β no explicit intermediate guidance β so long-horizon tasks suffer compounding errors. The paper frames this as an instruction-grounding gap and seeks a scalable post-training recipe to close it.
AtomVLA is "the first subtask-aware VLA framework integrated with a scalable offline post-training pipeline."
- Atomic-subtask decomposition β an LLM decomposes high-level demonstrations into fine-grained atomic subtasks (the intermediate guidance the SFT signal lacks).
- Latent world-model scoring β a pretrained predictive world model scores candidate action chunks against subtask goals in latent space, mitigating error accumulation and improving long-horizon robustness. (This is the reconstruction-free, deployable latent-WAM flavor β cf. Ο-0/DYNA-2 β used as an evaluator/critic, not an in-path generator.)
- Offline GRPO β the latent scoring enables Group Relative Policy Optimization without online physical-robot rollouts (the expensive part of RL-for-VLA β cf. DyGRO-VLA, RL for VLA).
- LIBERO 97.0%, LIBERO-PRO 48.0% avg success vs baselines; strong robustness under perturbations.
- Real-world on the Galaxea R1 Lite β broad applicability, especially long-horizon tasks.
- Datasets / checkpoints / code to be publicly released.
AtomVLA is a clean instance of IROS 2026's WAM trend (survey Β§5.1): the world model has moved from pixel generator to latent evaluator inside a VLA post-training loop. It combines three threads the wiki tracks β latent WAM (World Models), subtask/instruction grounding (Multi-Task VLA, AtomVLA's "instruction gap" β LangForce/DISC's language-shortcut problem), and offline RL for VLA (DyGRO-VLA). The offline-GRPO-via-latent-scoring move is the notable systems contribution: RL-quality post-training without robot rollouts.
Limitations (reviewer): results are LIBERO/LIBERO-PRO sim + one real platform; the world model's scoring fidelity is the ceiling (a bad latent critic misranks chunks); LLM subtask decomposition quality is an unmeasured dependency.
- Official program: IROS 2026 (paper #33) Β· survey: IROS 2026
- Related: World Models Β· VLA Hybrid Architectures Β· DyGRO-VLA Β· Multi-Task VLA Β· Ο-0
β Back to IROS 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)