Skip to content

IROS 2026 VLA RL

hwoo.han edited this page Sep 7, 2026 · 1 revision

IROS 2026 β€” VLA-RL: Masterful & General Manipulation with Scalable Reinforcement Learning

Venue: IROS 2026 (Pittsburgh) Β· paper #4707 Β· Tsinghua Univ. (Shenzhen IGS) Β· Nanyang Technological University (Lu, Guo, Zhang, Zhou, Jiang, Gao, Tang, Wang). Paper: arXiv 2505.18719 (May 2025) Β· HF. The scalable-RL datapoint of IROS 2026 β€” online RL that improves a pretrained autoregressive VLA by casting manipulation as a multi-turn conversation and supplying dense reward via a VLM process-reward model. Companions: RL for VLA Β· Multi-Task VLA Β· DyGRO-VLA Β· IROS 2026 survey.

VLA-RL framework β€” Rollout phase: N vectorized envs (curriculum-selected) run parallel rollouts of an OpenVLA+LoRA policy; each trajectory's sparse reward R is augmented with a dense Robotic Process Reward Model signal R^rprm from a fine-tuned VLM. Learning phase: GAE + a value/policy network update the policy via PPO over a replay buffer; actions are produced by OpenVLA (LoRA) β†’ action detokenizer β†’ (Ξ”x, Δθ, grip) (framework figure from Lu et al., arXiv 2505.18719, Β© the authors)

1. Problem

High-capacity VLAs imitate human demonstrations well, but limited state coverage in offline data causes failures out of distribution. An exploration-based method that improves from online data at test time can close this gap β€” but making online RL compatible with autoregressive VLAs (sparse rewards, huge action-token spaces, unstable/slow training) is the obstacle.

2. Method

VLA-RL is a systematic framework to online-RL-finetune pretrained autoregressive VLAs:

  • Trajectory-as-conversation β€” casts a manipulation trajectory as a multi-modal, multi-turn conversation, making RL optimization compatible with autoregressive VLAs.
  • Robotic Process Reward Model (RPRM) β€” a pretrained VLM fine-tuned as a process reward model, trained with pseudo reward labels from automatically extracted task segments, to densify sparse task rewards.
  • Systems for scale/stability β€” curriculum selection, GPU-balanced vectorized environments, batch decoding, and critic warmup; optimized via PPO with GAE.

3. Results

  • OpenVLA-7B surpasses the strongest finetuned baseline by +4.5% on 40 challenging LIBERO tasks, and matches commercial Ο€0-FAST.
  • Under a unified real-world protocol, VLA-RL lifts OpenVLA success 60% β†’ 90% within 3k interaction steps.
  • Keeps improving with more test-time optimization β€” an early sign of inference-scaling laws in robotics (more RL compute β†’ more skill).

4. Why it matters (RL / multi-task lens)

VLA-RL is IROS 2026's strongest evidence for the survey Β§3.2 "policy learning at scale" thread and a key entry in the Multi-Task VLA cluster C (optimization). It differs from DyGRO-VLA (which fights cross-task forgetting) by focusing on making online RL work at all for autoregressive VLAs at scale β€” the conversation reformulation + VLM process-reward are the enabling tricks. The inference-scaling observation (more RL budget keeps helping) is the notable forward-looking claim, echoing the LLM RL-scaling story. Contrast AtomVLA, which gets RL-quality post-training offline (WM critic, no rollouts) β€” VLA-RL pays for online rollouts but gets true exploration.

Limitations (reviewer): online RL needs many real/sim rollouts (3k steps real is cheap for RL but still a fleet cost); the RPRM's pseudo-reward quality bounds the signal; OpenVLA-7B + LIBERO is the main substrate; autoregressive-VLA-specific (not flow/diffusion heads).

5. Links

← Back to IROS 2026 survey Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally