-
Notifications
You must be signed in to change notification settings - Fork 0
CVPR 2026 MergeVLA
Venue: CVPR 2026 Β· arXiv: 2511.18810 (Nov 2025, rev. Mar 2026) Β· project: mergevla.github.io Authors: Yuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang, Zi Huang, Yadan Luo (University of Queensland) The anchor case study in Multi-Task VLA β compose per-task specialists into one generalist instead of jointly training.
VLA models fine-tune well on a single task/embodiment but degrade in multi-skill settings, and β the paper's focus β directly merging VLA experts trained on different tasks yields near-zero success. Two sources of non-mergeability:
- LoRA divergence in the VLM β "finetuning drives LoRA adapters in the VLM backbone toward divergent, task-specific directions beyond the capacity of existing merging methods to unify."
- Action-expert entanglement β "action experts develop inter-block dependencies through self-attention feedback, causing task information to spread across layers and preventing modular recombination."
Three architectural changes make specialists composable:
- Sparse task-masked LoRA β adapters are sparsely activated via task masks, retaining consistent parameters and reducing irreconcilable conflicts in the VLM.
- Cross-attention-only action expert β replaces self-attention with cross-attention-only blocks so each task's specialization stays localized and composable (no cross-block smearing).
- Test-time task router β when the task is unknown, adaptively selects the task mask + expert head from the initial observation, enabling unsupervised task inference.
Across LIBERO, LIBERO-Plus, RoboTwin, and a real SO101 robotic arm, MergeVLA achieves performance comparable to or exceeding individually finetuned experts β one merged checkpoint β N specialists β with robust generalization across tasks, embodiments, and environments. (The abstract reports relative parity rather than absolute per-benchmark numbers.)
MergeVLA reframes multi-task VLA as a merging problem and gives a transferable architectural lesson: self-attention in the action expert entangles tasks; cross-attention-only + task-masked LoRA keeps them composable. See the failure-mode taxonomy and the full solution landscape in Multi-Task VLA.
- arXiv: 2511.18810 Β· project: mergevla.github.io
- Review: Multi-Task VLA Β· VLA Architectures Β· VLMβAction Connection
β Back to Multi-Task VLA Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)