Skip to content

CoRL 2025 DSRL

Heungwoo edited this page Jun 1, 2026 · 2 revisions

DSRL β€” Steering Your Diffusion Policy with Latent Space Reinforcement Learning

Venue: CoRL 2025 (Award Finalist) Β· arXiv: 2506.15799 Authors: Andrew Wagenmaker, Mitsuhiko Nakamoto, Yunchu Zhang, Seohong Park, Waleed Yagoub, Anusha Nagabandi, Abhishek Gupta, Sergey Levine (UC Berkeley Β· University of Washington Β· Amazon) Category: RL for VLA / Diffusion Policies Trend tag: RL without log-probs

Approach diagram

flowchart LR
  Z[Initial noise z0] --> FROZEN[Frozen diffusion policy]
  FROZEN --> A[Action chunk]
  A --> R[Environment reward]
  R --> LEARN[Train a policy over z0<br/>RL action space = noise latent]
  LEARN --> Z
Loading

Problem

Diffusion / flow-matching policies don't expose action log-probabilities, so the standard RL toolkit (policy gradient, PPO) can't be applied out of the box. Prior workarounds either finetune the whole model (unstable) or approximate log-probs (noisy).

Method

DSRL (Diffusion Steering via Reinforcement Learning) makes a structural observation: the initial noise fed into a diffusion policy largely determines the sampled trajectory. Therefore you can freeze the diffusion policy and train a small RL actor that chooses the initial noise zβ‚€ given an observation. The noise latent becomes the RL action space β€” with real log-probs and cheap rollouts. Concretely, DSRL trains a lightweight off-policy actor-critic (SAC) over the latent-noise input, requires only black-box access to the BC policy (it never touches the base weights or gradients), and adds only a small set of trainable parameters (~500K).

Results

Post-hoc improvement of frozen diffusion policies across simulated benchmarks (OpenAI Gym, Robomimic, OGBench), single-task and multi-task (BridgeData V2) diffusion policies, and the pretrained generalist Ο€β‚€ policy (steered from its public DROID weights). DSRL reports substantially higher sample efficiency than direct-finetuning RL baselines (e.g., RLPD), and demonstrates real-world autonomous improvement on a physical Franka setup. CoRL 2025 Award Finalist.

Significance

The seed idea for the ICLR 2026 wave of RL-for-flow-matching papers. Every approach on the ICLR 2026 RL topic page can trace an ancestor to DSRL or its contemporaries:

Links

Related pages

← Back to CoRL-2025

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally