-
Notifications
You must be signed in to change notification settings - Fork 0
Review Multitask VLA
Scope: why a single Vision-Language-Action model degrades when asked to do many tasks at once, and the 2025β2026 research fixing it. Anchor case study: MergeVLA (CVPR 2026). Companions: VLA Architectures Β· VLMβAction Connection Β· Stellar VLA (continual) Β· RL for VLA.
Naively training (or merging) a single VLA over many tasks reliably underperforms per-task specialists. Six documented failure modes:
- Negative transfer / gradient conflict. Conflicting per-task gradients push shared weights toward compromises. HiMoE-VLA shows this directly β co-training a dense Ο0 across mismatched action spaces (joint-angle vs EE) lowers its score. It worsens under LoRA's low-rank capacity limit.
- Non-mergeability of specialists. Even merging separately-finetuned experts fails: MergeVLA finds naive merging gives near-zero success, because (a) LoRA adapters diverge in task-specific directions and (b) action-expert self-attention feedback spreads task information across blocks, preventing modular recombination.
- Multi-task conflict + instability at scale. As task count grows, joint optimization destabilizes. DyGRO-VLA documents that RL fine-tuning becomes unstable and less effective as tasks grow, and a t-SNE study shows single-task tuning drifts the task into an isolated cluster, away from the shared feature space.
- Catastrophic forgetting. Optimizing one task erodes others β DyGRO: LIBERO-Spatial RFT sharply drops unrelated LIBERO-Object success; the continual setting is worse (Stellar VLA, Pretrained VLAs Resist Forgetting).
- Routing confusion / skill fragmentation. In MoE policies, routing on low-level signals (e.g. diffusion noise level) fragments reusable skills across experts at skill boundaries, hurting transfer (SMoDP).
- Instruction collapse / wrong-task selection. With many tasks sharing scenes, the model latches onto visual shortcuts and ignores the language, executing the wrong task β the "Information Collapse" of LangForce and the "observation leakage" of DISC.
Root tension: a fixed-capacity shared backbone must both specialize (per-task precision) and generalize (cross-task competency) β an information-theoretic trade-off ("capability and robustness cannot both be free").
MergeVLA (arXiv 2511.18810; Fu et al., Univ. of Queensland) tackles the non-mergeability failure directly: compose per-task specialists into one generalist instead of jointly training.
Diagnosis β two barriers to merging:
- LoRA divergence β finetuning drives the VLM's LoRA adapters toward divergent task-specific directions beyond what merging can unify.
- Action-expert entanglement β self-attention feedback creates inter-block dependencies, so task information smears across layers and can't be recombined modularly.
Three fixes:
- Sparse task-masked LoRA β adapters are sparsely activated via task masks, keeping parameters consistent and reducing irreconcilable conflicts.
- Cross-attention-only action expert β replaces self-attention with cross-attention-only blocks, keeping each task's specialization localized and composable.
- Test-time task router β when the task is unknown, an unsupervised router picks the task mask + expert head from the initial observation.
Result: across LIBERO, LIBERO-Plus, RoboTwin, and a real SO101 arm, MergeVLA matches or exceeds individually finetuned experts β i.e. one merged checkpoint β N specialists. The architectural lesson generalizes: avoid self-attention-driven entanglement in the action expert if you want composability.
| Cluster | Idea | Exemplars |
|---|---|---|
| A. Model merging | train specialists, compose them into one checkpoint | MergeVLA (task-mask LoRA + cross-attn-only + test-time router) |
| B. Mixture-of-Experts (route per task/skill) | give each task/skill its own expert; shared backbone only consolidates | HiMoE-VLA (hierarchical AS-MoE/HB-MoE, 32 experts top-4), SMoDP (language-defined skill routing), SMP (orthogonal skill basis), HALO (3-expert MoT) |
| C. Optimization / gradient control | resolve gradient conflict during training | DyGRO-VLA (dynamic grouped residual optimization for cross-task RL), orthogonal gradient projection (2601.09684) |
| D. Skill decomposition / hierarchy | split a task into sub-skills, route per-phase | SMoDP, HALO (EM-CoT), Move-Then-Operate (phase experts), Ο0.5 hierarchy |
| E. Instruction grounding (right-task selection) | force the policy to actually follow language | DISC (decouple instruction from state), LangForce (penalize the vision shortcut), InstructVLA (instruction tuning) |
| F. Continual / knowledge-structured | add tasks over time without forgetting | Stellar VLA (Dirichlet-Process self-evolving knowledge space), Pretrained VLAs Resist Forgetting |
| G. Modular adapters / plug-ins | isolate task-specific capacity in small modules | DECO (plug-in tactile adapter), GuidedVLA (specialized attention heads), DexVLA plug-in expert |
- You already have strong per-task specialists? β merge them (A, MergeVLA) β cheapest path to one checkpoint, no joint retraining.
- Training one model on many tasks from scratch, heterogeneous action spaces/embodiments? β routed MoE (B, HiMoE-VLA) so mismatched tasks don't average out; add skill-semantic routing (SMoDP) to stop fragmentation.
- RL fine-tuning across tasks? β gradient/optimization control (C, DyGRO-VLA) to avoid the isolated-cluster forgetting.
- Long-horizon / compositional tasks? β skill decomposition (D) β route per sub-skill.
- Model executes the wrong task / ignores language? β instruction grounding (E, DISC/LangForce).
- Tasks arrive over time? β continual, knowledge-structured (F, Stellar VLA).
The unifying insight: every fix is a way to give each task its own parameters/pathway while sharing a common substrate β whether by masking (MergeVLA), routing (MoE), projecting gradients (DyGRO), or structuring a knowledge space (Stellar). Dense parameter sharing is exactly what causes the interference; structured, sparse specialization is the cure.
- Merge vs jointly-train vs route β which wins at matched compute? No head-to-head across MergeVLA (merge), HiMoE-VLA (MoE), and dense co-training on the same task suite.
- How many experts/masks before the router itself becomes the bottleneck? Test-time task inference (MergeVLA) and top-k gating (HiMoE) are unproven at hundreds of tasks.
- Does skill-semantic routing (SMoDP) generalize to unseen skill compositions, or only seen skills?
- Capabilityβrobustness bound β is the information-theoretic trade-off fundamental, or can structured specialization sidestep it?
- Cross-embodiment Γ multi-task β HiMoE handles action-space heterogeneity; combining with single-checkpoint multi-robot deployment is open.
- Anchor: MergeVLA (2511.18810, CVPR 2026, project)
- MoE / routing: HiMoE-VLA Β· SMoDP Β· SMP Β· HALO
- Optimization / continual: DyGRO-VLA Β· Stellar VLA Β· Pretrained VLAs Resist Forgetting
- IROS 2026 π: VLA-RL (scalable RL for masterful general manipulation) Β· LAR-MoE (latent-aligned routing MoE in imitation β cluster B) Β· AtomVLA (subtask grounding + offline GRPO β cluster C) Β· Responsibility-Induced Specialized Experts (MoE-VLA for humanoid loco-manip) β context: IROS 2026 survey
- Instruction grounding: DISC Β· LangForce Β· InstructVLA
- Adapters / decomposition: DECO Β· GuidedVLA Β· Move-Then-Operate
- Taxonomy: VLA Architectures Β· VLMβAction Connection
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)