Skip to content

Review Multitask VLA

hwoo.han edited this page Sep 7, 2026 · 4 revisions

In-Depth Review β€” Multi-Task VLA: why one policy struggles across tasks, and how to fix it

Scope: why a single Vision-Language-Action model degrades when asked to do many tasks at once, and the 2025–2026 research fixing it. Anchor case study: MergeVLA (CVPR 2026). Companions: VLA Architectures Β· VLM↔Action Connection Β· Stellar VLA (continual) Β· RL for VLA.


1. The problem β€” why one VLA β‰  many tasks

Naively training (or merging) a single VLA over many tasks reliably underperforms per-task specialists. Six documented failure modes:

  1. Negative transfer / gradient conflict. Conflicting per-task gradients push shared weights toward compromises. HiMoE-VLA shows this directly β€” co-training a dense Ο€0 across mismatched action spaces (joint-angle vs EE) lowers its score. It worsens under LoRA's low-rank capacity limit.
  2. Non-mergeability of specialists. Even merging separately-finetuned experts fails: MergeVLA finds naive merging gives near-zero success, because (a) LoRA adapters diverge in task-specific directions and (b) action-expert self-attention feedback spreads task information across blocks, preventing modular recombination.
  3. Multi-task conflict + instability at scale. As task count grows, joint optimization destabilizes. DyGRO-VLA documents that RL fine-tuning becomes unstable and less effective as tasks grow, and a t-SNE study shows single-task tuning drifts the task into an isolated cluster, away from the shared feature space.
  4. Catastrophic forgetting. Optimizing one task erodes others β€” DyGRO: LIBERO-Spatial RFT sharply drops unrelated LIBERO-Object success; the continual setting is worse (Stellar VLA, Pretrained VLAs Resist Forgetting).
  5. Routing confusion / skill fragmentation. In MoE policies, routing on low-level signals (e.g. diffusion noise level) fragments reusable skills across experts at skill boundaries, hurting transfer (SMoDP).
  6. Instruction collapse / wrong-task selection. With many tasks sharing scenes, the model latches onto visual shortcuts and ignores the language, executing the wrong task β€” the "Information Collapse" of LangForce and the "observation leakage" of DISC.

Root tension: a fixed-capacity shared backbone must both specialize (per-task precision) and generalize (cross-task competency) β€” an information-theoretic trade-off ("capability and robustness cannot both be free").


2. Anchor case study β€” MergeVLA (CVPR 2026)

MergeVLA (arXiv 2511.18810; Fu et al., Univ. of Queensland) tackles the non-mergeability failure directly: compose per-task specialists into one generalist instead of jointly training.

Diagnosis β€” two barriers to merging:

  • LoRA divergence β€” finetuning drives the VLM's LoRA adapters toward divergent task-specific directions beyond what merging can unify.
  • Action-expert entanglement β€” self-attention feedback creates inter-block dependencies, so task information smears across layers and can't be recombined modularly.

Three fixes:

  1. Sparse task-masked LoRA β€” adapters are sparsely activated via task masks, keeping parameters consistent and reducing irreconcilable conflicts.
  2. Cross-attention-only action expert β€” replaces self-attention with cross-attention-only blocks, keeping each task's specialization localized and composable.
  3. Test-time task router β€” when the task is unknown, an unsupervised router picks the task mask + expert head from the initial observation.

Result: across LIBERO, LIBERO-Plus, RoboTwin, and a real SO101 arm, MergeVLA matches or exceeds individually finetuned experts β€” i.e. one merged checkpoint β‰ˆ N specialists. The architectural lesson generalizes: avoid self-attention-driven entanglement in the action expert if you want composability.


3. The solution landscape

Cluster Idea Exemplars
A. Model merging train specialists, compose them into one checkpoint MergeVLA (task-mask LoRA + cross-attn-only + test-time router)
B. Mixture-of-Experts (route per task/skill) give each task/skill its own expert; shared backbone only consolidates HiMoE-VLA (hierarchical AS-MoE/HB-MoE, 32 experts top-4), SMoDP (language-defined skill routing), SMP (orthogonal skill basis), HALO (3-expert MoT)
C. Optimization / gradient control resolve gradient conflict during training DyGRO-VLA (dynamic grouped residual optimization for cross-task RL), orthogonal gradient projection (2601.09684)
D. Skill decomposition / hierarchy split a task into sub-skills, route per-phase SMoDP, HALO (EM-CoT), Move-Then-Operate (phase experts), Ο€0.5 hierarchy
E. Instruction grounding (right-task selection) force the policy to actually follow language DISC (decouple instruction from state), LangForce (penalize the vision shortcut), InstructVLA (instruction tuning)
F. Continual / knowledge-structured add tasks over time without forgetting Stellar VLA (Dirichlet-Process self-evolving knowledge space), Pretrained VLAs Resist Forgetting
G. Modular adapters / plug-ins isolate task-specific capacity in small modules DECO (plug-in tactile adapter), GuidedVLA (specialized attention heads), DexVLA plug-in expert

4. How the clusters relate β€” a decision guide

  • You already have strong per-task specialists? β†’ merge them (A, MergeVLA) β€” cheapest path to one checkpoint, no joint retraining.
  • Training one model on many tasks from scratch, heterogeneous action spaces/embodiments? β†’ routed MoE (B, HiMoE-VLA) so mismatched tasks don't average out; add skill-semantic routing (SMoDP) to stop fragmentation.
  • RL fine-tuning across tasks? β†’ gradient/optimization control (C, DyGRO-VLA) to avoid the isolated-cluster forgetting.
  • Long-horizon / compositional tasks? β†’ skill decomposition (D) β€” route per sub-skill.
  • Model executes the wrong task / ignores language? β†’ instruction grounding (E, DISC/LangForce).
  • Tasks arrive over time? β†’ continual, knowledge-structured (F, Stellar VLA).

The unifying insight: every fix is a way to give each task its own parameters/pathway while sharing a common substrate β€” whether by masking (MergeVLA), routing (MoE), projecting gradients (DyGRO), or structuring a knowledge space (Stellar). Dense parameter sharing is exactly what causes the interference; structured, sparse specialization is the cure.


5. Open questions

  1. Merge vs jointly-train vs route β€” which wins at matched compute? No head-to-head across MergeVLA (merge), HiMoE-VLA (MoE), and dense co-training on the same task suite.
  2. How many experts/masks before the router itself becomes the bottleneck? Test-time task inference (MergeVLA) and top-k gating (HiMoE) are unproven at hundreds of tasks.
  3. Does skill-semantic routing (SMoDP) generalize to unseen skill compositions, or only seen skills?
  4. Capability↔robustness bound β€” is the information-theoretic trade-off fundamental, or can structured specialization sidestep it?
  5. Cross-embodiment Γ— multi-task β€” HiMoE handles action-space heterogeneity; combining with single-checkpoint multi-robot deployment is open.

6. Links

← Back to Reviews Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally