-
Notifications
You must be signed in to change notification settings - Fork 0
CoRL 2025 Streaming Flow Policy
Venue: CoRL 2025 (Oral) Β· arXiv: 2505.21851 Authors: Sunshine Jiang, Xiaolin Fang, Nicholas Roy, TomΓ‘s Lozano-PΓ©rez, Leslie Pack Kaelbling, Siddharth Ancha (MIT) Also: Best Paper Nominee (3/21), ICRA 2025 Beyond Pick-and-Place Workshop Category: Flow / Diffusion Policies Trend tag: Trajectory streaming
flowchart LR
H[Observation history h] --> V[History-conditioned<br/>velocity field vΞΈ a,t|h]
V -- integrate from aβa_prev --> T["Flow time t = execution time"]
T --> A1[action a_t1]
T --> A2[action a_t2]
T --> A3[action a_t3]
A1 & A2 & A3 --> EXEC[Stream to robot on-the-fly]
Standard diffusion/flow policies sample an entire action chunk from noise in the trajectory space π^T before any action can execute, paying full denoising latency up front and producing discontinuities at chunk boundaries.
Streaming Flow Policy treats the action trajectory itself as the flow trajectory: it learns a velocity field directly in action space π and integrates it so that the flow-integration time coincides with execution time. Integration starts from a narrow Gaussian around the previous action, and each integration step emits the next action, which can be streamed to the robot on-the-fly during sampling (no "trajectory of trajectories", no chunk windowing).
-
History-conditioned velocity field
vΞΈ(a, t | h): inputs are the current actiona, normalized flow timet β [0,1], and observation historyh. -
Stabilizing conditional flow: each demonstration ΞΎ defines
v(a,t) = ΞΎΜ(t) β kΒ·(a β ΞΎ(t)), a feedback term with gainkthat contracts a thin Gaussian "tube" (varianceΟβΒ²Β·e^(β2kt)) around the demo, reducing distribution shift. -
Multimodality is preserved by training the marginal velocity over a mixture of these per-demonstration tubes
p*(a|t,h) = β« pΞΎ(a|t)Β·pπ(ΞΎ|h) dΞΎ, matching the per-timestep training distribution without sampling whole trajectories.
CoRL 2025 Oral. On Push-T (state), SFP reaches 95.1 / 96.0 (avg/max) vs Diffusion Policy 92.9 / 94.4 and flow-matching policy 80.6 / 82.6. On RoboMimic (state): Lift 100/100, Can 98.4/100 (DP 94.8/98.0), Square 78.0/84.0 (DP 77.2/84.0). Per-action latency β 3.5 ms vs DP-100-DDPM 40.2 ms and 10-step DDIM 4.4 ms; the streaming property lets sampling overlap with execution, avoiding chunk-boundary stalls and jerk.
Part of the CoRL 2025 trajectory-streaming trend alongside SAIL and DemoSpeedup. By collapsing flow time into execution time, it removes the chunk-boundary problem for flow-matching policies entirely rather than smoothing over it. Related to ICLR 2026's runtime-efficiency work (FASTER, OmniSAT, HyperVLA).
- arXiv: https://arxiv.org/abs/2505.21851
- Project page: https://streaming-flow-policy.github.io
β Back to CoRL-2025
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)