-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 PLD
Venue: ICLR 2026 Β· Project: wenlixiao.com/self-improve-VLA-PLD Category: RL for VLA Trend tag: Trend 3
flowchart LR
Step1[1. PROBE<br/>Distribution-aware data collection<br/>find failure cases of base VLA] --> Step2[2. LEARN<br/>Train lightweight residual actors<br/>off-policy RL SAC on probed failures<br/>BACKBONE FROZEN]
Step2 --> Step3[3. DISTILL<br/>Distill residual-corrected behavior<br/>back into base VLA]
Step3 --> Out[Self-improved VLA<br/>~99% LIBERO Β· 100% real Franka & YAM]
Step3 -. iterate .-> Step1
Standard RL fine-tuning of a pretrained VLA destabilizes the backbone β on-policy gradients pull the full VLA policy away from its pretrained distribution, causing catastrophic capability loss. Freezing the backbone and only training an action head doesn't help either β the information bottleneck is too tight. (Base policies evaluated are Ο0 (flow-matching action head) and OpenVLA (autoregressive action tokens); the paper does not quote a single fixed parameter count.)
Three stages:
- Probe β identify failure cases via distribution-aware data collection
- Learn β train lightweight residual actors on top of the frozen base VLA using sample-efficient off-policy RL (SAC, Gaussian policy) on probed failures; the residual outputs zero most of the time and only intervenes in failure states (only the residual updates online β backbone stable)
- Distill β distill the residual-corrected behavior back into the base VLA β self-improved model
Loop is repeatable.
- ~99% task success on LIBERO (near saturation)
- >50% performance gain on SimplerEnv
- 100% success on real-world Franka and YAM dexterous manipulation tasks
Among the strongest VLA RL numbers published.
Most compelling single demonstration that RL fine-tuning of VLAs is practical. Probe-Learn-Distill is repeatable and generalizes across platforms. LIBERO's saturation after PLD is the main trigger for the benchmark-generation work of RoboArena β and WorldGym.
- Project page: https://wenlixiao.com/self-improve-VLA-PLD
- OpenReview: https://openreview.net/forum?id=xqRD4LLY5M
- VLA-RFT (world-model alternative)
- Ο*0.6 + RECAP (flow-matching-specific alternative)
- RFS (residual-RL applied to dexterity)
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)