Embed paper architecture figures in Motus/HALO/BAGEL; add hybrid link to Home topic reviews - Motus: embed the paper's tri-expert architecture figure (Video Gen / Action / Understanding + Tri-modal Joint Attention), arXiv 2512.13030. - BAGEL: embed the MoT figure (Und/Gen experts + Multi-modal Self-Attention, dual encoders), arXiv 2505.14683. - HALO: embed Fig.1 (three-expert MoT) and Fig.2 (EM-CoT data pipeline). All figures credited to the authors alongside the schematic mermaids. - Home: add [[VLA Hybrid Architectures]] to the Topic reviews row. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add in-depth pages for HALO and BAGEL; add architecture diagram to Motus - Review-HALO (arXiv 2602.21157, ICML'26): three-expert MoT (AR understanding + diffusion visual-subgoal + flow-matching action), EM-CoT think→imagine→act, ~4.5B (Qwen2.5-1.5B×3); RoboTwin2.0 Easy 80.5% (+34.1 over pi0); full ablation analysis isolating the vision tower's value. Cross-linked with ICML-2026-HALO. - Review-BAGEL (arXiv 2505.14683, ByteDance): the base unified multimodal MoT recipe (understanding+generation, VAE+ViT dual encoders, separate QKV/FFN + shared attention, 7B/14B) that BagelVLA/HALO inherit; framed for VLA relevance. - Review-MOTUS: added mermaid architecture diagram (4 towers + scheduler). - Linked both from Review-VLA-Hybrid-Architectures core table/§8, ICML-2026-HALO, and the Reviews catalog. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>