Skip to content

CVPR 2026 MergeVLA

hwoo.han edited this page Aug 14, 2026 · 1 revision

MergeVLA β€” Cross-Skill Model Merging Toward a Generalist VLA

Venue: CVPR 2026 Β· arXiv: 2511.18810 (Nov 2025, rev. Mar 2026) Β· project: mergevla.github.io Authors: Yuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang, Zi Huang, Yadan Luo (University of Queensland) The anchor case study in Multi-Task VLA β€” compose per-task specialists into one generalist instead of jointly training.

Problem

VLA models fine-tune well on a single task/embodiment but degrade in multi-skill settings, and β€” the paper's focus β€” directly merging VLA experts trained on different tasks yields near-zero success. Two sources of non-mergeability:

  1. LoRA divergence in the VLM β€” "finetuning drives LoRA adapters in the VLM backbone toward divergent, task-specific directions beyond the capacity of existing merging methods to unify."
  2. Action-expert entanglement β€” "action experts develop inter-block dependencies through self-attention feedback, causing task information to spread across layers and preventing modular recombination."

Method

Three architectural changes make specialists composable:

  1. Sparse task-masked LoRA β€” adapters are sparsely activated via task masks, retaining consistent parameters and reducing irreconcilable conflicts in the VLM.
  2. Cross-attention-only action expert β€” replaces self-attention with cross-attention-only blocks so each task's specialization stays localized and composable (no cross-block smearing).
  3. Test-time task router β€” when the task is unknown, adaptively selects the task mask + expert head from the initial observation, enabling unsupervised task inference.

Results

Across LIBERO, LIBERO-Plus, RoboTwin, and a real SO101 robotic arm, MergeVLA achieves performance comparable to or exceeding individually finetuned experts β€” one merged checkpoint β‰ˆ N specialists β€” with robust generalization across tasks, embodiments, and environments. (The abstract reports relative parity rather than absolute per-benchmark numbers.)

Significance

MergeVLA reframes multi-task VLA as a merging problem and gives a transferable architectural lesson: self-attention in the action expert entangles tasks; cross-attention-only + task-masked LoRA keeps them composable. See the failure-mode taxonomy and the full solution landscape in Multi-Task VLA.

Links

← Back to Multi-Task VLA Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally