-
Notifications
You must be signed in to change notification settings - Fork 0
Review HiMoE VLA
Paper: "HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies" β arXiv 2512.05693 Β· ICLR 2026 Β· Fudan University & Microsoft Research Asia (Zhiying Du, Bei Liu, β¦ Yu-Gang Jiang) What it is: the MoE-routing answer to multi-task/cross-embodiment negative transfer β put a depth-wise mixture-of-experts inside the action module so heterogeneous tasks don't average out. Cluster B in Multi-Task VLA Β§3. Venue entry: ICLR-2026-HiMoE-VLA Β· companions X-VLA Β· XR-1 Β· Cross-Embodiment.
- Negative transfer is real and measurable. Co-training a dense Ο0 across mismatched action spaces (joint-angle vs end-effector) lowers its score β the paper's decisive datapoint that one dense backbone can't serve heterogeneous robots.
- Fix = depth-wise hierarchical MoE in the action module. Shallow AS-MoE layers specialize by action space; deeper HB-MoE layers absorb broader heterogeneity (embodiment kinematics, sensors); central dense blocks consolidate the shared cross-domain knowledge. Standard top-k gating (32 experts, top-4); no hand-built "arm vs hand" router.
- Two regularizers shape the experts: AS-Reg (contrastive β same-action-space tokens are positive pairs) sharpens action-space specialization; HB-Reg (route-frequency β expected-probability alignment) load-balances.
- SOTA across sim + real on a ~4B backbone: LIBERO 97.8%, CALVIN DβD 3.967, real xArm7 75.0%, real dual-arm ALOHA 63.7% β all above Ο0.
- It localizes the interference. The multi-task failure (Review-Multitask-VLA Β§1) is dense parameter sharing forcing compromises. HiMoE's answer is structural: give each action-space/embodiment its own feed-forward experts, and only consolidate in the central dense blocks β so specialization and sharing happen in different layers.
- Depth matters, not just sparsity. A single flat MoE (3.813) underperforms the hierarchy (4.012), and removing MoE entirely drops to 3.777 β evidence that where you separate (shallow=action-space, deep=embodiment) is the contribution, not merely "add experts."
- Complementary to the other cross-embodiment bets. It's the MoE-in-FFN counterpart to X-VLA (heterogeneity in the input prompt) and XR-1 (heterogeneity in the shared codebook) β three different loci for the same problem.
flowchart TB
In[Action-module input tokens] --> AS[Shallow β AS-MoE<br/>action-space specialization<br/>joint-angle vs end-effector Β· top-4/32]
AS --> HB[Deeper β HB-MoE<br/>embodiment / sensor heterogeneity Β· top-4/32]
HB --> D[Central dense Transformer blocks<br/>shared cross-domain knowledge]
D --> A[Action]
ASR[AS-Reg β contrastive on same-action-space experts] -.-> AS
HBR[HB-Reg β route-frequency load balance] -.-> HB
- Depth-wise hierarchy (not a category classifier): shallow layers handle the coarsest split (action space); deeper layers handle finer heterogeneity; dense blocks fuse.
- Routing: standard token-level top-k (32 experts, top-4) per MoE layer.
- AS-Reg: contrastive objective pulling together experts routed by same-action-space tokens β cleaner action-space specialization.
- HB-Reg: aligns empirical routing frequency with expected probability β even expert utilization (anti-collapse).
- Backbone: ~4B-param VLM.
| Benchmark | HiMoE-VLA | Ο0 | others |
|---|---|---|---|
| LIBERO (avg) | 97.8% | 94.2% | OpenVLA-OFT 97.1 Β· UniVLA 95.2 |
| CALVIN DβD (avg consecutive) | 3.967 | 3.758 | β |
| Real xArm7 (avg success) | 75.0% | 62.5% | β |
| Real ALOHA dual-arm | 63.7% | 54.2% | RDT-1B 47.5 |
Negative-transfer ablation (co-train CALVIN-ABC EEF + CALVIN-D joint): dense Ο0 degrades β0.259 (3.806β3.547) while HiMoE improves +0.186 (3.826β4.012). Remove-all-MoE β 3.777; flat MoE 3.813 < hierarchy 4.012.
Significance. HiMoE-VLA is the cleanest demonstration that hierarchical, depth-wise MoE converts cross-task/embodiment heterogeneity from a liability (negative transfer) into a gain β the routing-based cure in the multi-task fix landscape.
Limitations (authors' + reviewer's).
- All VLM-layer features are fused without selective weighting β a coarse consolidation.
- ~4B is modest β scale is constrained by available robotics data; behavior at larger scale / many more experts is untested.
- Not zero-shot to new robots β it demonstrates quick adaptation (e.g. 50k steps), not per-robot-fine-tune-free deployment (contrast single-checkpoint deployment).
- Router scaling β top-k gating at 32 experts is shown; hundreds of tasks/experts unproven.
Design lesson. Separate what conflicts (action-space in shallow experts, embodiment in deep experts) and share what transfers (central dense blocks) β depth-wise placement beats a flat expert pool.
- Paper: arXiv 2512.05693 Β· OpenReview Β· venue entry ICLR-2026-HiMoE-VLA
- Multi-task context: Multi-Task VLA (cluster B) Β· sibling optimizer approach DyGRO-VLA
- Cross-embodiment kin: X-VLA Β· XR-1 Β· Cross-Embodiment Β· VLA Architectures
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)