Skip to content

CoRL 2025 Streaming Flow Policy

Heungwoo edited this page Jun 1, 2026 · 2 revisions

Streaming Flow Policy

Venue: CoRL 2025 (Oral) Β· arXiv: 2505.21851 Authors: Sunshine Jiang, Xiaolin Fang, Nicholas Roy, TomΓ‘s Lozano-PΓ©rez, Leslie Pack Kaelbling, Siddharth Ancha (MIT) Also: Best Paper Nominee (3/21), ICRA 2025 Beyond Pick-and-Place Workshop Category: Flow / Diffusion Policies Trend tag: Trajectory streaming

Approach diagram

flowchart LR
  H[Observation history h] --> V[History-conditioned<br/>velocity field vΞΈ a,t|h]
  V -- integrate from aβ‰ˆa_prev --> T["Flow time t = execution time"]
  T --> A1[action a_t1]
  T --> A2[action a_t2]
  T --> A3[action a_t3]
  A1 & A2 & A3 --> EXEC[Stream to robot on-the-fly]
Loading

Problem

Standard diffusion/flow policies sample an entire action chunk from noise in the trajectory space π’œ^T before any action can execute, paying full denoising latency up front and producing discontinuities at chunk boundaries.

Method

Streaming Flow Policy treats the action trajectory itself as the flow trajectory: it learns a velocity field directly in action space π’œ and integrates it so that the flow-integration time coincides with execution time. Integration starts from a narrow Gaussian around the previous action, and each integration step emits the next action, which can be streamed to the robot on-the-fly during sampling (no "trajectory of trajectories", no chunk windowing).

  • History-conditioned velocity field vΞΈ(a, t | h): inputs are the current action a, normalized flow time t ∈ [0,1], and observation history h.
  • Stabilizing conditional flow: each demonstration ΞΎ defines v(a,t) = ΞΎΜ‡(t) βˆ’ kΒ·(a βˆ’ ΞΎ(t)), a feedback term with gain k that contracts a thin Gaussian "tube" (variance Οƒβ‚€Β²Β·e^(βˆ’2kt)) around the demo, reducing distribution shift.
  • Multimodality is preserved by training the marginal velocity over a mixture of these per-demonstration tubes p*(a|t,h) = ∫ pΞΎ(a|t)Β·pπ’Ÿ(ΞΎ|h) dΞΎ, matching the per-timestep training distribution without sampling whole trajectories.

Results

CoRL 2025 Oral. On Push-T (state), SFP reaches 95.1 / 96.0 (avg/max) vs Diffusion Policy 92.9 / 94.4 and flow-matching policy 80.6 / 82.6. On RoboMimic (state): Lift 100/100, Can 98.4/100 (DP 94.8/98.0), Square 78.0/84.0 (DP 77.2/84.0). Per-action latency β‰ˆ 3.5 ms vs DP-100-DDPM 40.2 ms and 10-step DDIM 4.4 ms; the streaming property lets sampling overlap with execution, avoiding chunk-boundary stalls and jerk.

Significance

Part of the CoRL 2025 trajectory-streaming trend alongside SAIL and DemoSpeedup. By collapsing flow time into execution time, it removes the chunk-boundary problem for flow-matching policies entirely rather than smoothing over it. Related to ICLR 2026's runtime-efficiency work (FASTER, OmniSAT, HyperVLA).

Links

Related pages

← Back to CoRL-2025

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally