Skip to content

Review HALO

hwoo.han edited this page Aug 11, 2026 · 2 revisions

In-Depth Review β€” HALO: a three-expert VLA for Embodied Multimodal Chain-of-Thought

Model: HALO β€” unified three-expert Mixture-of-Transformers VLA (think β†’ imagine β†’ act) Β· HKUST Β· EPFL Β· Sun Yat-sen University (Quanxin Shou, Fangqi Zhu, …; corresp. Song Guo) Paper: arXiv 2602.21157 Β· ICML 2026 (poster, venue page) Β· survey entry: ICML-2026-HALO One-liner: the clearest academic instance of the three-expert MoT with a dedicated visual tower β€” and the one that ablates why the vision tower earns its place.

Part of the WAM+VLA hybrid family: VLA Hybrid Architectures (comparison + design guide) Β· siblings Motus Β· BagelVLA Β· BAGEL Β· DYNA-2.


1. TL;DR

  1. Three experts, one shared self-attention. HALO is a Mixture-of-Transformers with a Multimodal-Understanding expert (autoregressive text), a Visual-Generation expert (diffusion / flow-matching β†’ subgoal image), and an Action expert (flow-matching β†’ action chunk). The experts keep independent parameters but share the self-attention; special tokens (<visual_start>, <action_start>) route the active modality.
  2. Embodied Multimodal Chain-of-Thought (EM-CoT). It reasons like a person: textual reasoning r β†’ visual subgoal Γ΄ β†’ action a, each conditioned on the last (Eqs. 1–3). "Think in words, imagine in pixels, then act."
  3. Small backbone, big gains. Each expert is a Qwen2.5-1.5B (~4.5B total). On RoboTwin 2.0 (50 tasks): 80.5% Easy / 26.4% Hard, beating Ο€0 by +34.1 / +10.1 points; on real Cobot Mobile ALOHA it leads Ο€0/Ο€0.5 on all four tasks.
  4. The ablations are the contribution. Removing the visual-generation data alone drops Hard success >50%; no pre-training β†’ 0% on Hard. Both the textual chain and the visual subgoal are needed for the OOD robustness.

2. Why it matters

  • It isolates the value of the third (vision) tower. Where Motus and DYNA-2 argue for a vision tower at scale, HALO ablates it at small scale: dropping visual-generation data costs >50% on Hard tasks, and the visual-subgoal branch is separately necessary. This is the cleanest published evidence that a dedicated visual-foresight tower does real work, not just regularization.
  • EM-CoT unifies two previously separate CoT lines. Prior work adds either textual chain-of-thought or visual subgoal prediction; HALO fuses both into one sequential reasoning process inside one model β€” a structured answer to long-horizon / OOD manipulation.
  • It runs the vision tower in-path (as a subgoal), yet stays a small model. Unlike full-video WAMs (3–4Γ— latency), HALO's vision tower emits a single subgoal image, keeping the "imagine" step affordable β€” the mid-point of the granularity axis in Review-VLA-Hybrid-Architectures Β§4(a).

3. Architecture

HALO's unified Mixture-of-Transformers β€” three experts (Multimodal Understanding Β· Visual Generation Β· Action Prediction) with independent parameters sharing one self-attention (Figure 1 from Shou et al., 2026, Β© the authors)

flowchart LR
  L[instruction + obs history] --> U[Understanding expert Β· AR<br/>Qwen2.5-1.5B Β· ViT+SigLIP2/NaViT]
  U -->|reasoning r| V[Visual-Generation expert Β· diffusion<br/>FLUX VAE, subgoal image Γ΄]
  V -->|subgoal Γ΄| A[Action expert Β· flow-matching<br/>action chunk a]
  U <-. shared self-attention .-> V
  V <-. shared self-attention .-> A
  U <-. shared self-attention .-> A
  A ==> OUT[action]
Loading
  • Understanding expert (AR). Qwen2.5-1.5B (28 layers, 12 heads, hidden 1536). Vision for understanding: ViT + SigLIP2 (384Β²β†’980Β²) with NaViT for native aspect ratios.
  • Visual-Generation expert (diffusion). Predicts the subgoal image via flow-matching (MSE); pixels through a frozen FLUX VAE (8Γ— downsample, 16 latent channels).
  • Action expert (flow-matching, L₁). Low-dim continuous actions via linear projection.
  • Fusion. Independent parameter sets per expert, shared self-attention; a switching mechanism (special tokens) sets the active modality. Attention masking: causal for AR text, bidirectional within a frame / causal across frames, noise tokens masked from clean context.
  • EM-CoT (Eqs. 1–3): r ~ P(Β·|l,o) β†’ Γ΄_{t+h} ~ P(Β·|l,o,r) β†’ a_{t:t+m} ~ Ο€(Β·|l,o,r,Γ΄).

Training β€” two stages.

  • Stage 1 β€” Versatile Pre-training (90k steps): VQA (LLaVA-NeXT-779k, CE) + Visual Generation (OXE + SSv2 video, flow-MSE) + Action (OXE, L₁). Loss L = 0.25Β·L_CE + 0.5Β·L_MSE + L_L1.
  • Stage 2 β€” EM-CoT-Augmented Fine-tuning (110k sim / 80k real): an automated 3-phase EM-CoT pipeline β€” (1) low-level actions β†’ motion primitives by rule-matching, (2) Qwen3-VL adds dense textual reasoning + subtask decomposition, (3) each subtask's terminal frame = visual subgoal. Corpus: 2,500 sim demos + 320 real demos, co-trained with general VQA to prevent forgetting. Loss L = L_r + L_Γ΄ + L_a.

HALO's automated EM-CoT data pipeline β€” low-level actions β†’ motion primitives (rule-based), VLM-augmented dense textual reasoning, and terminal-frame visual subgoals (Figure 2 from Shou et al., 2026, Β© the authors)


4. Results (paper-reported)

RoboTwin 2.0 (50 tasks, 100 trials each; baselines from the official leaderboard):

Setting HALO Ο€0 best baseline (RDT-1B)
Easy 80.5% 46.4% (+34.1) 34.5%
Hard 26.4% 16.3% (+10.1) 13.7%

Hard-task highlights: Stack-Blocks-Three 37% (Ο€0 0%); Shake-Bottle 73% (Ο€0 60%).

Real world β€” Cobot Mobile ALOHA, 4 tasks Γ— 50 trials:

Task HALO Ο€0 Ο€0.5
Sweep buttons 98% 84% 82%
Bimanual cup nesting 92% 64% 68%
Screwdriver handover 88% 72% 76%
Lemon β†’ drawer 94% 70% 74%

Ablations (the core evidence):

Pre-training config Easy Hard
Full (Vision + Text + Action) 75.3% 21.2%
w/o Visual-generation data 58.2% 10.5%
w/o Vision + Text VQA 42.9% 3.9%
w/o pre-training 32.4% 0%

EM-CoT: w/o textual reasoning 77.8 / 18.3; w/o visual subgoal 76.1 / 22.5; both (HALO) 80.5 / 26.4. Even HALO-w/o-EM-CoT beats Ο€0 by +28.9 on Easy.


5. Significance & limitations

Significance. HALO is the strongest controlled case that the three-expert (vision-as-tower) MoT is more than a parameter dump: each tower's data measurably drives success, EM-CoT compounds them, and the win is largest on Hard / OOD tasks β€” exactly where monolithic VLAs fail. It also shows the recipe works at ~4.5B, not just 8–14B.

Limitations.

  1. arXiv/ICML-poster; unreplicated externally. Numbers are the authors' own (leaderboard-anchored baselines).
  2. Vision tower runs in-path. A subgoal image is cheaper than a video rollout but still an extra generative step per chunk; no latency table is reported.
  3. Heavy reliance on pre-training diversity (0% Hard without it) β€” the recipe is data-recipe-sensitive, not just architecture-driven.
  4. Small backbone caps language breadth (Qwen2.5-1.5B experts); the EM-CoT text is task-scoped, not open-domain reasoning.
  5. No explicit limitations section; the failure envelope is inferred from ablations.

6. Links

← Back to Reviews Β· ICML-2026 Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally