-
Notifications
You must be signed in to change notification settings - Fork 0
ICML 2026 HALO
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: HKUST, EPFL, Sun Yat-sen University (Quanxin Shou, Fangqi Zhu, et al.; corresp. Song Guo) Traction (2026-06): 2 citations (arXiv)

Vision-Language-Action (VLA) models perform well on manipulation but struggle in long-horizon or out-of-distribution scenarios because they lack explicit mechanisms for multimodal reasoning and for anticipating how the world will evolve under action. Recent work adds either textual chain-of-thought or visual subgoal prediction, but no method offers a unified human-like reasoning framework that jointly does textual reasoning, visual foresight, and action prediction.
HALO is a unified VLA model that performs embodied multimodal chain-of-thought (EM-CoT) reasoning as a sequential process: textual task reasoning β visual subgoal prediction (fine-grained guidance) β EM-CoT-augmented action prediction.
- Unified architecture. Inspired by BAGEL's harmonization of multimodal tasks, HALO uses a Mixture-of-Transformers (MoT) with three specialized experts β Multimodal Understanding, Visual Generation, and Action Prediction β that keep independent parameter sets but interact through a shared self-attention mechanism, enabling rich cross-modal collaboration while a switching mechanism controls the active modality.
- EM-CoT data pipeline. An automatic three-phase pipeline synthesizes EM-CoT training data at scale: (1) translate continuous low-level actions into high-level motion primitives via rule-based matching, (2) use a large VLM (Qwen3-VL) to add dense textual reasoning β task narratives and subtask decomposition, and (3) designate each subtask's terminal frame as its visual subgoal, giving sparse supervision that lowers learning difficulty.
- Two-stage training recipe. Stage 1 β Versatile Pre-training over a heterogeneous mix of VQA, Visual Generation, and Action Prediction data to build a generalist foundation. Stage 2 β EM-CoT-Augmented Fine-tuning on the synthesized reasoning/subgoal-aligned data to elicit structured multimodal reasoning.

Evaluated on RoboTwin 2.0 (50 tasks) against Οβ, RDT, and Diffusion Policy, with baseline numbers taken from the official RoboTwin 2.0 leaderboard.
- Simulation. HALO reaches 80.46% on Easy and 26.44% on Hard tasks, surpassing Οβ by 34.1% and 10.1% respectively. The relative gaps over Οβ (73.5% Easy, 62.0% Hard) are large especially on Hard tasks, indicating stronger OOD robustness.
- Ablations. Even without the explicit reasoning chain, HALO-w/o-EM-CoT already beats the strongest baseline (Οβ) by +28.92 points; all training-recipe and EM-CoT components further improve success rate.
- Real-world. Across four basic tasks (tool-use sweeping, bimanual cup nesting, inter-arm screwdriver handover, drawer placement) plus a generalization setting with visual distractions, lighting/background changes, and novel objects, HALO shows strong generalization under aggressive unseen randomization, with accurate textual reasoning and subgoal image generation.
HALO unifies textual reasoning, visual foresight, and action into a single MoT-based VLA, rather than bolting one reasoning modality onto a policy. The shared-attention three-expert design plus an automated EM-CoT data pipeline lets the model "think in words, imagine in pixels, then act" β yielding both higher success rates and notably better out-of-distribution robustness on a demanding bimanual benchmark.
- In-depth review: HALO β in-depth (full architecture, EM-CoT equations, ablation analysis, hybrid-family positioning)
- Hybrid-family comparison: VLA Hybrid Architectures
- arXiv: 2602.21157
- ICML 2026: https://icml.cc/virtual/2026/poster/61922
β Back to ICML-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)