-
Notifications
You must be signed in to change notification settings - Fork 0
ICML 2026 Move Then Operate
Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation β A dual-expert VLA that splits coarse relocation from contact-critical interaction
Venue: ICML 2026 (Poster) Category: VLA Architecture Traction (2026-06): 0 citations (arXiv)

Monolithic VLA policies use a single network to handle two very different behavioral regimes: coarse relocation ("move" β transporting the end-effector through free space) and contact-critical interaction ("operate" β fine, high-precision manipulation against an object). Conflating these heterogeneous dynamics creates optimization interference, where gradients for free-space transit and gradients for delicate contact conflict. Move-Then-Operate asks whether explicitly disentangling these phases is a more effective and data-efficient inductive bias, aligning the policy with human motor patterns.
The architecture is a dual-expert policy built on a flow-matching (conditional flow matching) VLA. A shared vision-language encoder feeds two expert heads, E_move and E_operate, which share the base architecture but keep disjoint parameters so that conflicting gradient updates between coarse transit and fine manipulation are isolated.
- A latent variable z β {Move, Operate} performs hard routing: a single expert defines the entire vector field, and the selection stays invariant across the whole flow-integration interval Ο β [0,1] for each generated action chunk.
- Routing is done per action chunk by a learnable phase selector trained with supervised routing learning.
- Phase-aware auto-labeling: an MLLM β³ takes the demonstration video and instruction and predicts a hierarchical schedule of N consecutive subtasks, each decomposed into atomic phases of type {Move, Operate}. Topological constraints (e.g., decomposition depth β€ 2 per subtask) enforce physically plausible labels conditioned on lightweight cues such as end-effector velocity.

On the RoboTwin2 benchmark with a constrained budget of 50 demonstrations per task, Move-Then-Operate reaches a 68.9% average success rate, outperforming the monolithic Οβ baseline by +24.1%. Gains concentrate on high-precision tasks: +55% on Click Bell and +8% on Press Stapler over the next-best method, and long-horizon Place Cans improves from 64% β 79%.
The method is also data- and compute-efficient: trained on only 50 tasks Γ 50 demos (100,000 steps, no per-task fine-tuning), it matches or exceeds Οβ.β * and GO-1* trained on ~10Γ more data, and reaches peak performance in 40% fewer training steps. An ablation confirms the routing is load-bearing β replacing the learned router with Random Selection collapses success from 68.88% to 25.63%, and an adversarial Reversal Selection drops it to 8.88%, showing the two experts specialize in distinct, incompatible behaviors.
Move-Then-Operate demonstrates that a simple structural prior β separating "move" from "operate" with disjoint experts and a learned per-chunk router β substantially mitigates optimization interference in VLAs, yielding strong precision gains while being markedly more data- and training-efficient than scaling data. The MLLM auto-labeling pipeline makes the phase supervision cheap to obtain.
- arXiv: 2604.23620
- ICML 2026: https://icml.cc/virtual/2026/poster/65496
β Back to ICML-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)