Skip to content

ICLR 2026 Flow Matching PG

Heungwoo edited this page Jun 1, 2026 · 4 revisions

Flow Matching Policy Gradients (FPO) β€” On-Policy RL with Flow-Matching Actors

Venue: ICLR 2026 Β· OpenReview: eoEmoKoQpJ Category: RL for VLA β€” Policy parameterization Trend tag: Flow-matching policies / on-policy RL

Approach diagram

flowchart LR
  Pol[Flow-matching policy vΜ‚_ΞΈ<br/>denoising MLP / VLA head] --> Roll[Rollouts: any sampler<br/>10-step Euler / stochastic / det.]
  Roll --> Adv[GAE advantage Γ‚_t]
  Roll --> MC["Store N_mc<br/>(Ο„_i, Ξ΅_i) pairs per (o_t,a_t)"]
  Adv --> Ratio["FPO ratio:<br/>rΜ‚ = exp(L_CFM,ΞΈ_old βˆ’ L_CFM,ΞΈ)"]
  MC --> Ratio
  Ratio --> Clip["PPO-clip surrogate<br/>min(rΜ‚Γ‚, clip(rΜ‚)Γ‚)"]
  Clip --> Pol
Loading

Problem

Standard policy-gradient methods (PPO) parameterise actions as diagonal Gaussians, which struggle with multimodal action distributions and under-conditioned tasks (e.g. humanoid root-only goal conditioning). Flow / diffusion policies capture multimodality well but lack tractable likelihoods, making them awkward to plug into the PPO-clip ratio. Existing on-policy RL methods for diffusion (DDPO, DPPO, Flow-GRPO) treat each denoising step as its own MDP transition, which (i) multiplies horizon length by 10–50, (ii) treats initial noise as part of observation, increasing problem dimensionality, and (iii) restricts to stochastic samplers.

Detailed Method

FPO is a drop-in replacement for the PPO likelihood ratio that uses the conditional flow matching (CFM) loss as a likelihood proxy.

Background: Conditional Flow Matching

Given clean action a_t and noise Ξ΅ ∼ N(0,I), interpolate at flow timestep Ο„ ∈ [0,1]:

a_t^Ο„ = (1-Ο„) a_t + Ο„ Ξ΅ (OT schedule)

Train velocity field vΜ‚_ΞΈ to predict the conditional flow u(a_t^Ο„, Ο„ | a_t) = a_t βˆ’ Ξ΅:

L_CFM,ΞΈ = E_{Ο„, q(a), p_Ο„(a^Ο„|a)} [ β€– vΜ‚_ΞΈ(a^Ο„, Ο„) βˆ’ (a βˆ’ Ξ΅) β€–Β² ]

FPO ratio (Eq. 6, 16)

Replace the PPO log-likelihood ratio with a CFM-loss-based proxy:

rΜ‚_FPO(ΞΈ) = exp( LΜ‚_CFM,ΞΈ_old(a_t; o_t) βˆ’ LΜ‚_CFM,ΞΈ(a_t; o_t) )

where each LΜ‚ is a Monte-Carlo estimate over N_mc draws of (Ο„_i, Ξ΅_i):

LΜ‚_CFM,ΞΈ(a_t; o_t) = (1/N_mc) Ξ£_i β€– vΜ‚_ΞΈ(a_t^{Ο„_i}, Ο„_i; o_t) βˆ’ (a_t βˆ’ Ξ΅_i) β€–Β²

The (Ο„_i, Ξ΅_i) pairs are frozen between the ΞΈ_old and ΞΈ evaluations so that the difference cancels noise.

Surrogate objective (Eq. 5)

Plug rΜ‚_FPO into PPO-clip:

max_ΞΈ E_{a_tβˆΌΟ€_old} [ min( rΜ‚_FPO Γ‚_t, clip(rΜ‚_FPO, 1βˆ’Ξ΅_clip, 1+Ξ΅_clip) Γ‚_t ) ]

Why this is sound (Sec. 3.3)

Kingma & Gao (2023) show the diffusion-loss-with-constant-weight equals -ELBO + const; thus

r_FPO(ΞΈ) = exp(ELBO_ΞΈ - ELBO_ΞΈ_old) = (Ο€_ΞΈ/Ο€_ΞΈ_old) Β· exp(D_KL_ΞΈ_old βˆ’ D_KL_ΞΈ)

The first factor is the true likelihood ratio; the second factor tightens the variational bound. Maximising the FPO ratio simultaneously increases the modelled likelihood of high-advantage actions and tightens the ELBO.

Bias analysis

The Monte-Carlo single-sample estimator rΜ‚_FPO^(Ο„,Ξ΅) is an upward biased estimate of r_FPO (Jensen). However the gradient is unbiased (Eq. 19-20):

βˆ‡_ΞΈ rΜ‚_FPO = βˆ’ rΜ‚_FPO Β· βˆ‡_ΞΈ β„“_ΞΈ(Ο„, Ξ΅) E[ βˆ’βˆ‡_ΞΈ β„“_ΞΈ ] = βˆ‡_ΞΈ ELBO_ΞΈ

So even with N_mc = 1, FPO produces directionally unbiased gradients. Empirically, more samples help but a single sample already beats Gaussian PPO.

Algorithm 1

While not converged:
  Collect rollouts using any sampler; compute Γ‚_t with GAE
  For each action store N_mc (Ο„_i, Ξ΅_i) pairs and β„“_ΞΈ(Ο„_i, Ξ΅_i)
  ΞΈ_old ← ΞΈ
  For each epoch:
    For each minibatch (o_t, a_t, {(Ο„_i,Ξ΅_i)}):
      Compute β„“_ΞΈ(Ο„_i, Ξ΅_i)
      rΜ‚_ΞΈ = exp( -(1/N_mc) Ξ£_i (β„“_ΞΈ - β„“_ΞΈ_old) )
      L_FPO = min(rΜ‚Γ‚, clip(rΜ‚, 1Β±Ξ΅)Γ‚)
      ΞΈ ← Optimizer(ΞΈ, βˆ‡L_FPO)
  Update value head as in standard PPO

Hyperparameters (MuJoCo Playground)

  • Optimiser: Adam.
  • 60 M total environment steps, batch size 1024, 16 updates per batch.
  • 10 sampling steps for FPO and DPPO.
  • Learning rate: 3 Γ— 10⁻⁴ for FPO and DPPO.
  • Clip Ξ΅ swept ∈ {0.01, 0.05, 0.1, 0.2, 0.3}; final Ξ΅ = 0.05 for FPO.
  • For DPPO: per-step Gaussian noise Οƒ_t swept ∈ {0.01, 0.05, 0.1}; Οƒ_t = 0.05, Ξ΅ = 0.2 used.
  • N_mc = 8 in main experiments; ablations at 1 and 4.
  • Ξ΅-CFM (compute CFM loss on noise prediction Ξ΅Μ‚) used over u-CFM (CFM on velocity) for scale invariance.

Comprehensive Results

MuJoCo Playground (10 DM Control Suite tasks, 5 seeds, 60M steps)

Table 1 reports the average evaluation reward across MuJoCo tasks:

Method Avg Reward
Gaussian PPO 667.8 Β± 66.0
Gaussian PPO† (default HPs) 577.2 Β± 74.4
DPPO 652.5 Β± 83.7
FPO‑ 759.3 Β± 45.3
FPO, 1 (Ο„,Ξ΅) 691.6 Β± 50.3
FPO, 4 (Ο„,Ξ΅) 731.2 Β± 58.2
FPO, u-MSE (velocity) 664.6 Β± 48.5
FPO, Ξ΅_clip = 0.1 623.3 Β± 76.3
FPO, Ξ΅_clip = 0.2 526.4 Β± 76.8

‑ = 8 (Ο„,Ξ΅) pairs, Ξ΅-MSE, Ξ΅_clip = 0.05.

FPO wins on 8 of 10 Playground tasks (visualised in Figures 2 and 3 against Gaussian PPO and DPPO over BallInCup, FingerSpin, FingerTurnEasy/Hard, FishSwim, PointMass, ReacherEasy/Hard, CartpoleBalance, CheetahRun).

Humanoid Control (Isaac Gym, SMPL humanoid, 24 actuated joints Γ— 6 DoF)

Goal-conditioned MoCap tracking on AMASS, evaluated by success rate (joint distance ≀ 0.5 m), alive duration, and global MPJPE:

Method Goal Success ↑ Alive ↑ MPJPE ↓
Gaussian PPO All joints 98.7 % 200.46 31.62
FPO All joints 96.4 % 198.00 41.98
Gaussian PPO Root + Hands 46.5 % 142.50 97.65
FPO Root + Hands 70.6 % 171.32 62.91
Gaussian PPO Root only 29.8 % 114.06 123.70
FPO Root only 54.3 % 152.90 73.55

When sufficient conditioning is provided (all joints), Gaussian PPO is on par. As the conditioning becomes sparser (root or root+hands only), Gaussian PPO collapses while FPO retains 54–71 % success β€” a ~24-point gain. Prior methods solving sparse-goal humanoid control require a teacher β†’ student distillation pipeline; FPO learns end-to-end.

Beyond MoCap tracking, the authors also show FPO trains a humanoid that walks across procedurally generated rough terrain (Figure 4c).

GridWorld (25Γ—25 with two goal cells)

Designed to probe multimodality. FPO learns a bimodal action distribution at saddle-point states (where multiple optima exist), visible in the denoising flow visualisation (Figure 1). Gaussian PPO consistently picks the nearest goal with low diversity.

Ablation Studies

  • N_mc sweep (Table 1). N_mc = 1 β†’ 691.6, N_mc = 4 β†’ 731.2, N_mc = 8 β†’ 759.3. More samples help but FPO with a single sample still beats Gaussian PPO (667.8) and DPPO (652.5).
  • Ξ΅-CFM vs u-CFM. Ξ΅-MSE (predict noise) gives 759.3; u-MSE (predict velocity) drops to 664.6. The authors hypothesise Ξ΅ is invariant to action scale, which makes a single Ξ΅_clip transferable across tasks.
  • Clipping Ξ΅. Ξ΅ = 0.05 best (759.3); Ξ΅ = 0.1 β†’ 623.3; Ξ΅ = 0.2 β†’ 526.4. Tight clipping critical, similar to but tighter than Gaussian PPO.
  • DPPO denoising-MDP comparison (Sec. 3.5). DPPO multiplies horizon by 10–50 (one PPO step per denoising step) and is restricted to stochastic samplers; FPO is sampler-agnostic at both train and test.
  • Goal-conditioning sweep on humanoid. Full β†’ root+hands β†’ root reveals widening FPO advantage as conditioning sparsens.

Limitations

The authors do not include a separate "Limitations" section, but the Discussion surfaces:

  • Compute cost. "The training and deployment of flow-based policies is generally more computationally intensive than for corresponding Gaussian policies" (explicit Discussion statement).
  • Missing PPO machinery. FPO "lacks established machinery such as KL divergence estimation for adaptive learning rates and entropy regularization" β€” features standard PPO tooling provides for Gaussian policies.
  • No real-robot deployment β€” all experiments are in simulation (MuJoCo Playground, Isaac Gym, GridWorld). The rough-terrain locomotion result is described as showing "potential for sim-to-real transfer," but no hardware results are reported.
  • Image-diffusion fine-tuning unstable. The authors explored applying FPO to fine-tune a pretrained image diffusion model with RL and "found this setting to be unstable in practice," attributing the instability to the broader difficulty of RL on image generation rather than to FPO itself. (The paper does not evaluate or name large diffusion VLAs such as Ο€0 or GR00T-N1.)
  • MC-bias trade-off. Single-sample rΜ‚ is upward-biased; gradients are directionally unbiased but variance is higher. The authors use N_mc = 8 as the main-experiment default.

The paper's stated future direction is applying FPO where flow-based policies are already pretrained β€” e.g. behavior-cloned diffusion policies in robotics β€” where its PPO-clip compatibility and simplicity may help fine-tuning with task reward.

Significance & Positioning

FPO is the simplest known on-policy RL recipe for flow / diffusion policies that does not require approximating likelihoods or unrolling the sampler as an MDP. Three properties make it especially valuable for the VLA agenda:

  1. Sampler-agnostic. The same flow actor can be trained with any deterministic or stochastic integrator, and inference can use any number of steps. This decouples training compute from inference latency β€” important for VLAs that may want fast 1-step inference but stable many-step training.
  2. PPO-clip drop-in. Existing PPO codebases need only swap the likelihood ratio. The authors implemented FPO on top of Brax PPO with minimal changes.
  3. Under-conditioning gain. The 24-point success-rate gap on root-only humanoid control is the most direct evidence yet that flow policies' expressivity matters for under-specified tasks β€” exactly the regime where VLAs often operate (sparse language goal, multimodal answer).

Within the 2026 RL-for-VLA cluster:

  • VLA-RFT also targets RL fine-tuning of flow VLAs but inside a learned world model; FPO is model-free.
  • SimpleVLA-RL offers a Gaussian PPO recipe for VLA RL fine-tuning; FPO is the flow-native counterpart.
  • Guided Flow Policy is the offline-RL counterpart that adds Q-guidance to flow training; FPO complements it with online policy gradients.
  • Flow-To-Policy uses flows as planners rather than policies; FPO targets policies directly.
  • ReinFlow (NeurIPS 2025) explored a similar territory but with Gaussian-step formulation; FPO replaces the Gaussian-step trick with a CFM-loss ratio.
  • Compared to flow-VLA architectures themselves (Unified Diffusion VLA, dVLA, Ο€0 family), FPO provides the missing post-training tool: online RL without giving up the flow generative structure.

The paper is also one of the cleanest theoretical contributions of the cycle β€” the equivalence between FPO ratio and ELBO-ratio (Eq. 12) provides a justification that previous diffusion-RL recipes lacked.

Links

Related pages

← Back to ICLR-2026 Β· Topic: RL

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally