-
Notifications
You must be signed in to change notification settings - Fork 0
ICML 2026 OMP
Venue: ICML 2026 (Poster) Category: Diffusion-Flow Policy Affiliations: Han Fang, Yize Huang, Yuheng Zhao, Paul Weng, Xiao Li, Yutong Ban Traction (2026-06): 4 citations (arXiv)

Generative robot policies face a hard trade-off: diffusion models (DP3, NFE=10) achieve high success but suffer ~132 ms latency on Adroit, while flow-based methods (FlowPolicy, AdaFlow) reach single-step inference at the cost of architectural complexity and over-constrained training that hurts generalization. MeanFlow and its first robotics adaptation MP1 enable NFE=1 inference (6.8 ms, 19x faster than DP3) but have two flaws: the model predicts instantaneous velocity yet is trained on interval-averaged velocities (a mismatch that degrades trajectory accuracy), and the required Jacobian-Vector Product (JVP) operator consumes heavy GPU memory.
OMP improves MeanFlow-based policies along two axes. (1) Directional Alignment. The authors analyze the geometry of the MSE loss in high-dimensional velocity regression and show that the directional gradient term scales as 2Β·ΟΒ·Ο*Β·sin Ξ± β i.e., the gradient that corrects direction vanishes as the target velocity magnitude Ο*β0. This is exactly the regime of high-precision tasks (peg insert, thread-in-hole), explaining why standard flow models stall on fine manipulation. OMP adds a lightweight Cosine Loss that directly aligns the direction of the predicted interval-averaged velocity with the true mean velocity, restoring directional gradient signal independent of magnitude. (2) DDE for the JVP operator. OMP replaces the exact JVP with a Differential Derivation Equation (DDE) that approximates the Jacobian-Vector Product via finite differences, sharply reducing GPU memory for complex tasks at the cost of small approximation error. A Dispersive Loss term is retained for latent-feature discrimination.

Evaluated on 3 Adroit tasks and 34 Meta-World tasks (10 expert demos each, RTX 4090, three seeds). OMP reaches 82.3% Β± 1.7% average success, vs. MP1's 78.9% Β± 2.1% β a 3.4% gain over MP1 and 10.7% over FlowPolicy. Because MP1 is already near-optimal on the 21 Meta-World "easy" tasks (88.2%), the headline average understates OMP; on harder splits OMP improves Meta-World Medium by 9.4%, Hard by 4.4%, and Very-Hard by 10.6%. Training curves show OMP converges faster and more stably than baselines. Ablations confirm removing the directional alignment (Cosine Loss) causes a significant drop across all Adroit and Meta-World tasks; swapping JVP for DDE trades a small accuracy drop for reduced GPU memory.
OMP gives a clean theoretical account of why single-step flow policies fail on fine manipulation β the vanishing directional gradient β and fixes it with a near-free cosine term, while the DDE trick makes MeanFlow-style NFE=1 policies practical on memory-limited hardware. It pushes the real-time, single-step generative-policy frontier that matters for high-frequency robot control.
- arXiv: 2512.19347
- ICML 2026: https://icml.cc/virtual/2026/poster/66693
β Back to ICML-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)