Skip to content

ICLR 2026 HiMoE VLA

hwoo.han edited this page Aug 14, 2026 · 3 revisions

HiMoE-VLA β€” Hierarchical Mixture-of-Experts for Generalist VLA

Venue: ICLR 2026 (arXiv:2512.05693, Dec 2025) Authors: Zhiying Du, Bei Liu, Yaobo Liang, Yichao Shen, Haidong Cao, Xiangyu Zheng, Zhiyuan Feng, Zuxuan Wu, Jiaolong Yang, Yu-Gang Jiang (Fudan University & Microsoft Research Asia) Category: VLA Architecture β€” Cross-Embodiment Trend tag: Trend 4

Approach diagram

The hierarchy is depth-wise across the action module, not a category classifier. Each MoE layer uses standard top-k token routing (32 experts, top-4); there is no explicit "arm vs hand" router.

flowchart TB
  In[Action-module input tokens] --> AS[Shallow layers: AS-MoE<br/>action-space specialization<br/>joint-angle vs end-effector]
  AS --> HB[Deeper layers: HB-MoE<br/>balances broader heterogeneity<br/>embodiment / sensors]
  HB --> D[Central dense Transformer blocks<br/>shared cross-domain knowledge]
  D --> A[Action]
Loading

Problem

A single dense backbone trained jointly on heterogeneous robots suffers from negative transfer: different action spaces (joint-angle vs. end-effector control), embodiment kinematics, sensor configurations, and control frequencies push the shared weights toward compromises that hurt individual robots. The paper directly demonstrates this β€” co-training a dense Ο€0 across mismatched action spaces lowers its score.

Method

Hierarchical (depth-wise) Mixture-of-Experts in the action module: shallow AS-MoE layers specialize per action space (joint-angle vs. end-effector control); deeper HB-MoE layers balance broader heterogeneity (embodiment kinematics, sensor configs); central dense Transformer blocks consolidate the heterogeneous signals into shared representations for cross-domain transfer. Each MoE layer routes tokens with standard top-k gating (32 experts, top-4).

Two regularizers shape specialization:

  • AS-Reg β€” a contrastive objective treating experts assigned to the same action-space token as positive pairs, sharpening action-space specialization.
  • HB-Reg β€” aligns empirical routing frequency with expected routing probability, spreading heterogeneous inputs evenly across experts (load balancing).

The VLM backbone is ~4B params.

Results

State-of-the-art on simulation and real robots. LIBERO: 97.8% overall avg (vs UniVLA 95.2%, OpenVLA-OFT 97.1%, π0 94.2%). CALVIN D→D: 3.967 avg consecutive tasks (vs π0 3.758). Real xArm7: 75.0% avg success (π0 62.5%). Real ALOHA (dual-arm): 63.7% (π0 54.2%, RDT-1B 47.5%).

Decisive ablation on negative transfer: co-training CALVIN-ABC (EEF) + CALVIN-D (joint), Ο€0 degrades βˆ’0.259 (3.806β†’3.547) while HiMoE improves +0.186 (3.826β†’4.012). Removing all MoE layers drops the model to 3.777 vs 4.012; a single flat MoE (3.813) underperforms the hierarchy, confirming depth-wise separation matters.

Limitations (per authors): features from all VLM layers are fused without selective weighting; model scale (~4B) is modest, constrained by available robotics data. Note: the method demonstrates quick adaptation / fine-tuning to new robots (e.g. 50k steps), not zero per-robot fine-tuning.

Significance

The MoE-based counterpart to X-VLA (soft prompts) and XR-1 (UVMC). All three are bets on where in the architecture embodiment heterogeneity should be handled β€” HiMoE places it in the feed-forward layers, X-VLA in the input prompt, XR-1 in the tokenizer / shared codebook.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally