Skip to content

ICLR 2026 PLD

Heungwoo edited this page Jun 1, 2026 · 3 revisions

PLD β€” Probe, Learn, Distill (Self-Improving VLA via Residual RL)

Venue: ICLR 2026 Β· Project: wenlixiao.com/self-improve-VLA-PLD Category: RL for VLA Trend tag: Trend 3

Approach diagram

flowchart LR
  Step1[1. PROBE<br/>Distribution-aware data collection<br/>find failure cases of base VLA] --> Step2[2. LEARN<br/>Train lightweight residual actors<br/>off-policy RL SAC on probed failures<br/>BACKBONE FROZEN]
  Step2 --> Step3[3. DISTILL<br/>Distill residual-corrected behavior<br/>back into base VLA]
  Step3 --> Out[Self-improved VLA<br/>~99% LIBERO Β· 100% real Franka & YAM]
  Step3 -. iterate .-> Step1
Loading

Problem

Standard RL fine-tuning of a pretrained VLA destabilizes the backbone β€” on-policy gradients pull the full VLA policy away from its pretrained distribution, causing catastrophic capability loss. Freezing the backbone and only training an action head doesn't help either β€” the information bottleneck is too tight. (Base policies evaluated are Ο€0 (flow-matching action head) and OpenVLA (autoregressive action tokens); the paper does not quote a single fixed parameter count.)

Method

Three stages:

  1. Probe β€” identify failure cases via distribution-aware data collection
  2. Learn β€” train lightweight residual actors on top of the frozen base VLA using sample-efficient off-policy RL (SAC, Gaussian policy) on probed failures; the residual outputs zero most of the time and only intervenes in failure states (only the residual updates online β†’ backbone stable)
  3. Distill β€” distill the residual-corrected behavior back into the base VLA β†’ self-improved model

Loop is repeatable.

Results

  • ~99% task success on LIBERO (near saturation)
  • >50% performance gain on SimplerEnv
  • 100% success on real-world Franka and YAM dexterous manipulation tasks

Among the strongest VLA RL numbers published.

Significance

Most compelling single demonstration that RL fine-tuning of VLAs is practical. Probe-Learn-Distill is repeatable and generalizes across platforms. LIBERO's saturation after PLD is the main trigger for the benchmark-generation work of RoboArena ∞ and WorldGym.

Links

Related pages

← Back to ICLR-2026 Β· Topic: RL

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally