-
Notifications
You must be signed in to change notification settings - Fork 0
ICML 2026 VLAW
VLAW: Iterative Co-Improvement of VLA Policy and World Model β bootstrapping a real-world simulator from policy rollouts
Venue: ICML 2026 (Poster) Category: World Model Affiliations: Yanjiang Guo, Tony Lee, Lucy Xiaoyang Shi, Jianyu Chen, Percy Liang, Chelsea Finn (Stanford / Tsinghua) Traction (2026-06): 18 citations (arXiv)

Improving VLA models via online interaction is bottlenecked by the cost of real-world policy rollouts. A natural fix is to use a learned simulator β an action-conditioned video generation model β to manufacture extra rollout data. But existing world models lack the physical fidelity needed for policy improvement: trained predominantly on demonstration datasets that under-cover diverse physical interactions (especially failure cases), they fail to model the small but critical contact dynamics of object manipulation.
VLAW is a simple iterative co-improvement loop between a VLA policy and an action-conditioned video world model. Real-world rollout data improves the world model's fidelity, and the improved world model then generates synthetic rollouts that improve the policy:
-
World Model Learning with Real Roll-outs: the policy is rolled out
Ktimes in the real world to form a datasetD_real, with each trajectory assigned a sparse success/failure rewardr β {0,1}at robot reset. BecauseD_realcontains both successes and failures, it counters two known pathologies β over-optimism (training data dominated by successful demos) and limited physical fidelity on contact-rich/deformable dynamics. -
World model initialization: VLAW starts from a pretrained Ctrl-World diffusion-based world model (trained on the full DROID dataset) and fine-tunes it on
D_realwith the standard diffusion denoising objective. - Iterative policy improvement: the grounded world model generates large-scale synthetic rollouts used to improve the VLA policy; the paper relates this loop to regularized reinforcement learning (Appendix A).

"39.2% absolute success rate improvement over the base policy". On the DROID platform (Franka Panda + Robotiq gripper, two third-person + one wrist camera) across five contact-rich task categories (stacking, opening a book, and others), VLAW improves a state-of-the-art VLA model by 39.2% absolute success rate over the base policy, with 11.6% of that coming specifically from training on the world-model-generated synthetic rollouts. World-model quality metrics (PSNR/SSIM/LPIPS/FID/FVD plus an event confusion matrix) confirm that adding expert and online rollouts sharply improves fidelity over the pretrained Ctrl-World baseline (e.g., FVD dropping from 225.13 toward ~100 after expert-rollout fine-tuning).
VLAW shows that a video world model can be turned into a useful, physically grounded data engine for VLA post-training using only a small amount of real interaction β and that including failure rollouts is key to fidelity. This offers a scalable, sample-efficient alternative to pure real-world RL for VLA improvement on contact-rich tasks.
- arXiv: 2602.12063
- ICML 2026: https://icml.cc/virtual/2026/poster/66169
β Back to ICML-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)