-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 Flow Matching PG
Venue: ICLR 2026 Β· OpenReview: eoEmoKoQpJ Category: RL for VLA β Policy parameterization Trend tag: Flow-matching policies / on-policy RL
flowchart LR
Pol[Flow-matching policy vΜ_ΞΈ<br/>denoising MLP / VLA head] --> Roll[Rollouts: any sampler<br/>10-step Euler / stochastic / det.]
Roll --> Adv[GAE advantage Γ_t]
Roll --> MC["Store N_mc<br/>(Ο_i, Ξ΅_i) pairs per (o_t,a_t)"]
Adv --> Ratio["FPO ratio:<br/>rΜ = exp(L_CFM,ΞΈ_old β L_CFM,ΞΈ)"]
MC --> Ratio
Ratio --> Clip["PPO-clip surrogate<br/>min(rΜΓ, clip(rΜ)Γ)"]
Clip --> Pol
Standard policy-gradient methods (PPO) parameterise actions as diagonal Gaussians, which struggle with multimodal action distributions and under-conditioned tasks (e.g. humanoid root-only goal conditioning). Flow / diffusion policies capture multimodality well but lack tractable likelihoods, making them awkward to plug into the PPO-clip ratio. Existing on-policy RL methods for diffusion (DDPO, DPPO, Flow-GRPO) treat each denoising step as its own MDP transition, which (i) multiplies horizon length by 10β50, (ii) treats initial noise as part of observation, increasing problem dimensionality, and (iii) restricts to stochastic samplers.
FPO is a drop-in replacement for the PPO likelihood ratio that uses the conditional flow matching (CFM) loss as a likelihood proxy.
Given clean action a_t and noise Ξ΅ βΌ N(0,I), interpolate at flow timestep Ο β [0,1]:
a_t^Ο = (1-Ο) a_t + Ο Ξ΅ (OT schedule)
Train velocity field vΜ_ΞΈ to predict the conditional flow u(a_t^Ο, Ο | a_t) = a_t β Ξ΅:
L_CFM,ΞΈ = E_{Ο, q(a), p_Ο(a^Ο|a)} [ β vΜ_ΞΈ(a^Ο, Ο) β (a β Ξ΅) βΒ² ]
Replace the PPO log-likelihood ratio with a CFM-loss-based proxy:
rΜ_FPO(ΞΈ) = exp( LΜ_CFM,ΞΈ_old(a_t; o_t) β LΜ_CFM,ΞΈ(a_t; o_t) )
where each LΜ is a Monte-Carlo estimate over N_mc draws of (Ο_i, Ξ΅_i):
LΜ_CFM,ΞΈ(a_t; o_t) = (1/N_mc) Ξ£_i β vΜ_ΞΈ(a_t^{Ο_i}, Ο_i; o_t) β (a_t β Ξ΅_i) βΒ²
The (Ο_i, Ξ΅_i) pairs are frozen between the ΞΈ_old and ΞΈ evaluations so that the difference cancels noise.
Plug rΜ_FPO into PPO-clip:
max_ΞΈ E_{a_tβΌΟ_old} [ min( rΜ_FPO Γ_t, clip(rΜ_FPO, 1βΞ΅_clip, 1+Ξ΅_clip) Γ_t ) ]
Kingma & Gao (2023) show the diffusion-loss-with-constant-weight equals -ELBO + const; thus
r_FPO(ΞΈ) = exp(ELBO_ΞΈ - ELBO_ΞΈ_old) = (Ο_ΞΈ/Ο_ΞΈ_old) Β· exp(D_KL_ΞΈ_old β D_KL_ΞΈ)
The first factor is the true likelihood ratio; the second factor tightens the variational bound. Maximising the FPO ratio simultaneously increases the modelled likelihood of high-advantage actions and tightens the ELBO.
The Monte-Carlo single-sample estimator rΜ_FPO^(Ο,Ξ΅) is an upward biased estimate of r_FPO (Jensen). However the gradient is unbiased (Eq. 19-20):
β_ΞΈ rΜ_FPO = β rΜ_FPO Β· β_ΞΈ β_ΞΈ(Ο, Ξ΅) E[ ββ_ΞΈ β_ΞΈ ] = β_ΞΈ ELBO_ΞΈ
So even with N_mc = 1, FPO produces directionally unbiased gradients. Empirically, more samples help but a single sample already beats Gaussian PPO.
While not converged:
Collect rollouts using any sampler; compute Γ_t with GAE
For each action store N_mc (Ο_i, Ξ΅_i) pairs and β_ΞΈ(Ο_i, Ξ΅_i)
ΞΈ_old β ΞΈ
For each epoch:
For each minibatch (o_t, a_t, {(Ο_i,Ξ΅_i)}):
Compute β_ΞΈ(Ο_i, Ξ΅_i)
rΜ_ΞΈ = exp( -(1/N_mc) Ξ£_i (β_ΞΈ - β_ΞΈ_old) )
L_FPO = min(rΜΓ, clip(rΜ, 1Β±Ξ΅)Γ)
ΞΈ β Optimizer(ΞΈ, βL_FPO)
Update value head as in standard PPO
- Optimiser: Adam.
- 60 M total environment steps, batch size 1024, 16 updates per batch.
- 10 sampling steps for FPO and DPPO.
- Learning rate: 3 Γ 10β»β΄ for FPO and DPPO.
- Clip Ξ΅ swept β {0.01, 0.05, 0.1, 0.2, 0.3}; final Ξ΅ = 0.05 for FPO.
- For DPPO: per-step Gaussian noise Ο_t swept β {0.01, 0.05, 0.1}; Ο_t = 0.05, Ξ΅ = 0.2 used.
- N_mc = 8 in main experiments; ablations at 1 and 4.
- Ξ΅-CFM (compute CFM loss on noise prediction Ξ΅Μ) used over u-CFM (CFM on velocity) for scale invariance.
Table 1 reports the average evaluation reward across MuJoCo tasks:
| Method | Avg Reward |
|---|---|
| Gaussian PPO | 667.8 Β± 66.0 |
| Gaussian PPOβ (default HPs) | 577.2 Β± 74.4 |
| DPPO | 652.5 Β± 83.7 |
| FPOβ‘ | 759.3 Β± 45.3 |
| FPO, 1 (Ο,Ξ΅) | 691.6 Β± 50.3 |
| FPO, 4 (Ο,Ξ΅) | 731.2 Β± 58.2 |
| FPO, u-MSE (velocity) | 664.6 Β± 48.5 |
| FPO, Ξ΅_clip = 0.1 | 623.3 Β± 76.3 |
| FPO, Ξ΅_clip = 0.2 | 526.4 Β± 76.8 |
β‘ = 8 (Ο,Ξ΅) pairs, Ξ΅-MSE, Ξ΅_clip = 0.05.
FPO wins on 8 of 10 Playground tasks (visualised in Figures 2 and 3 against Gaussian PPO and DPPO over BallInCup, FingerSpin, FingerTurnEasy/Hard, FishSwim, PointMass, ReacherEasy/Hard, CartpoleBalance, CheetahRun).
Goal-conditioned MoCap tracking on AMASS, evaluated by success rate (joint distance β€ 0.5 m), alive duration, and global MPJPE:
| Method | Goal | Success β | Alive β | MPJPE β |
|---|---|---|---|---|
| Gaussian PPO | All joints | 98.7 % | 200.46 | 31.62 |
| FPO | All joints | 96.4 % | 198.00 | 41.98 |
| Gaussian PPO | Root + Hands | 46.5 % | 142.50 | 97.65 |
| FPO | Root + Hands | 70.6 % | 171.32 | 62.91 |
| Gaussian PPO | Root only | 29.8 % | 114.06 | 123.70 |
| FPO | Root only | 54.3 % | 152.90 | 73.55 |
When sufficient conditioning is provided (all joints), Gaussian PPO is on par. As the conditioning becomes sparser (root or root+hands only), Gaussian PPO collapses while FPO retains 54β71 % success β a ~24-point gain. Prior methods solving sparse-goal humanoid control require a teacher β student distillation pipeline; FPO learns end-to-end.
Beyond MoCap tracking, the authors also show FPO trains a humanoid that walks across procedurally generated rough terrain (Figure 4c).
Designed to probe multimodality. FPO learns a bimodal action distribution at saddle-point states (where multiple optima exist), visible in the denoising flow visualisation (Figure 1). Gaussian PPO consistently picks the nearest goal with low diversity.
- N_mc sweep (Table 1). N_mc = 1 β 691.6, N_mc = 4 β 731.2, N_mc = 8 β 759.3. More samples help but FPO with a single sample still beats Gaussian PPO (667.8) and DPPO (652.5).
- Ξ΅-CFM vs u-CFM. Ξ΅-MSE (predict noise) gives 759.3; u-MSE (predict velocity) drops to 664.6. The authors hypothesise Ξ΅ is invariant to action scale, which makes a single Ξ΅_clip transferable across tasks.
- Clipping Ξ΅. Ξ΅ = 0.05 best (759.3); Ξ΅ = 0.1 β 623.3; Ξ΅ = 0.2 β 526.4. Tight clipping critical, similar to but tighter than Gaussian PPO.
- DPPO denoising-MDP comparison (Sec. 3.5). DPPO multiplies horizon by 10β50 (one PPO step per denoising step) and is restricted to stochastic samplers; FPO is sampler-agnostic at both train and test.
- Goal-conditioning sweep on humanoid. Full β root+hands β root reveals widening FPO advantage as conditioning sparsens.
The authors do not include a separate "Limitations" section, but the Discussion surfaces:
- Compute cost. "The training and deployment of flow-based policies is generally more computationally intensive than for corresponding Gaussian policies" (explicit Discussion statement).
- Missing PPO machinery. FPO "lacks established machinery such as KL divergence estimation for adaptive learning rates and entropy regularization" β features standard PPO tooling provides for Gaussian policies.
- No real-robot deployment β all experiments are in simulation (MuJoCo Playground, Isaac Gym, GridWorld). The rough-terrain locomotion result is described as showing "potential for sim-to-real transfer," but no hardware results are reported.
- Image-diffusion fine-tuning unstable. The authors explored applying FPO to fine-tune a pretrained image diffusion model with RL and "found this setting to be unstable in practice," attributing the instability to the broader difficulty of RL on image generation rather than to FPO itself. (The paper does not evaluate or name large diffusion VLAs such as Ο0 or GR00T-N1.)
- MC-bias trade-off. Single-sample rΜ is upward-biased; gradients are directionally unbiased but variance is higher. The authors use N_mc = 8 as the main-experiment default.
The paper's stated future direction is applying FPO where flow-based policies are already pretrained β e.g. behavior-cloned diffusion policies in robotics β where its PPO-clip compatibility and simplicity may help fine-tuning with task reward.
FPO is the simplest known on-policy RL recipe for flow / diffusion policies that does not require approximating likelihoods or unrolling the sampler as an MDP. Three properties make it especially valuable for the VLA agenda:
- Sampler-agnostic. The same flow actor can be trained with any deterministic or stochastic integrator, and inference can use any number of steps. This decouples training compute from inference latency β important for VLAs that may want fast 1-step inference but stable many-step training.
- PPO-clip drop-in. Existing PPO codebases need only swap the likelihood ratio. The authors implemented FPO on top of Brax PPO with minimal changes.
- Under-conditioning gain. The 24-point success-rate gap on root-only humanoid control is the most direct evidence yet that flow policies' expressivity matters for under-specified tasks β exactly the regime where VLAs often operate (sparse language goal, multimodal answer).
Within the 2026 RL-for-VLA cluster:
- VLA-RFT also targets RL fine-tuning of flow VLAs but inside a learned world model; FPO is model-free.
- SimpleVLA-RL offers a Gaussian PPO recipe for VLA RL fine-tuning; FPO is the flow-native counterpart.
- Guided Flow Policy is the offline-RL counterpart that adds Q-guidance to flow training; FPO complements it with online policy gradients.
- Flow-To-Policy uses flows as planners rather than policies; FPO targets policies directly.
- ReinFlow (NeurIPS 2025) explored a similar territory but with Gaussian-step formulation; FPO replaces the Gaussian-step trick with a CFM-loss ratio.
- Compared to flow-VLA architectures themselves (Unified Diffusion VLA, dVLA, Ο0 family), FPO provides the missing post-training tool: online RL without giving up the flow generative structure.
The paper is also one of the cleanest theoretical contributions of the cycle β the equivalence between FPO ratio and ELBO-ratio (Eq. 12) provides a justification that previous diffusion-RL recipes lacked.
- OpenReview: https://openreview.net/forum?id=eoEmoKoQpJ
- Guided Flow Policy β offline-RL counterpart
- Flow-To-Policy β flow-as-planner alternative
- VLA-RFT β RL fine-tuning in world model
- SimpleVLA-RL
- ReinFlow
- Unified Diffusion VLA
- dVLA
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)