-
Notifications
You must be signed in to change notification settings - Fork 0
IROS 2026 VLA RL
Venue: IROS 2026 (Pittsburgh) Β· paper #4707 Β· Tsinghua Univ. (Shenzhen IGS) Β· Nanyang Technological University (Lu, Guo, Zhang, Zhou, Jiang, Gao, Tang, Wang). Paper: arXiv 2505.18719 (May 2025) Β· HF. The scalable-RL datapoint of IROS 2026 β online RL that improves a pretrained autoregressive VLA by casting manipulation as a multi-turn conversation and supplying dense reward via a VLM process-reward model. Companions: RL for VLA Β· Multi-Task VLA Β· DyGRO-VLA Β· IROS 2026 survey.

High-capacity VLAs imitate human demonstrations well, but limited state coverage in offline data causes failures out of distribution. An exploration-based method that improves from online data at test time can close this gap β but making online RL compatible with autoregressive VLAs (sparse rewards, huge action-token spaces, unstable/slow training) is the obstacle.
VLA-RL is a systematic framework to online-RL-finetune pretrained autoregressive VLAs:
- Trajectory-as-conversation β casts a manipulation trajectory as a multi-modal, multi-turn conversation, making RL optimization compatible with autoregressive VLAs.
- Robotic Process Reward Model (RPRM) β a pretrained VLM fine-tuned as a process reward model, trained with pseudo reward labels from automatically extracted task segments, to densify sparse task rewards.
- Systems for scale/stability β curriculum selection, GPU-balanced vectorized environments, batch decoding, and critic warmup; optimized via PPO with GAE.
- OpenVLA-7B surpasses the strongest finetuned baseline by +4.5% on 40 challenging LIBERO tasks, and matches commercial Ο0-FAST.
- Under a unified real-world protocol, VLA-RL lifts OpenVLA success 60% β 90% within 3k interaction steps.
- Keeps improving with more test-time optimization β an early sign of inference-scaling laws in robotics (more RL compute β more skill).
VLA-RL is IROS 2026's strongest evidence for the survey Β§3.2 "policy learning at scale" thread and a key entry in the Multi-Task VLA cluster C (optimization). It differs from DyGRO-VLA (which fights cross-task forgetting) by focusing on making online RL work at all for autoregressive VLAs at scale β the conversation reformulation + VLM process-reward are the enabling tricks. The inference-scaling observation (more RL budget keeps helping) is the notable forward-looking claim, echoing the LLM RL-scaling story. Contrast AtomVLA, which gets RL-quality post-training offline (WM critic, no rollouts) β VLA-RL pays for online rollouts but gets true exploration.
Limitations (reviewer): online RL needs many real/sim rollouts (3k steps real is cheap for RL but still a fleet cost); the RPRM's pseudo-reward quality bounds the signal; OpenVLA-7B + LIBERO is the main substrate; autoregressive-VLA-specific (not flow/diffusion heads).
- Paper: arXiv 2505.18719 Β· HF Β· official program: IROS 2026 (#4707)
- Related: RL for VLA Β· Multi-Task VLA Β· DyGRO-VLA Β· AtomVLA Β· IROS 2026 survey
β Back to IROS 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)