Skip to content

Review BAGEL

hwoo.han edited this page Aug 11, 2026 · 2 revisions

In-Depth Review β€” BAGEL: the unified multimodal MoT recipe robot VLAs inherit

Model: BAGEL-7B-MoT β€” open unified multimodal understanding+generation model Β· ByteDance-Seed Paper: "Emerging Properties in Unified Multimodal Pretraining" β€” arXiv 2505.14683 Β· HF ByteDance-Seed/BAGEL-7B-MoT Β· blog Why it's in this wiki: BAGEL is not a robot model β€” but its Mixture-of-Transformers recipe is the template that BagelVLA, HALO, and the whole three-expert vision-tower family build on. This page documents the base recipe so the robot pages can point to one source.

Family: VLA Hybrid Architectures Β· BagelVLA Β· HALO Β· Motus.


1. TL;DR

  1. One model, understanding + generation, via Mixture-of-Transformer-Experts. BAGEL runs two experts β€” multimodal understanding and visual generation β€” with separate QKV projectors and FFNs per expert but shared attention layers. 7B active / 14B total, initialized from Qwen2.5-7B-Instruct.
  2. Dual visual encoders β€” the key idea robot VLAs copy. A VAE supplies pixel-level features (for generation) and a ViT supplies semantic features (for understanding), so one model both sees and draws without the two objectives fighting.
  3. Trillions of interleaved tokens. Pre-trained on interleaved text Β· image Β· video Β· web data; capability appears in stages (understanding+generation early β†’ basic editing β†’ complex editing).
  4. Emergent "world-modeling." Beyond image gen/edit/understand, it shows future-frame prediction, free-form manipulation, multiview synthesis, and world navigation β€” the very capabilities a WAM needs, which is why robot hybrids adopt its backbone.

2. Why it matters for VLA

  • It is the base recipe of the three-expert MoT. HALO states it is "inspired by BAGEL's harmonization of multimodal tasks"; BagelVLA literally initializes from BAGEL and adds a third (action) expert. The pattern β€” per-expert QKV/FFN + shared self-attention, plus VAE+ViT dual vision encoders β€” is what lets robot models bolt on a video/action tower without wrecking language grounding (Review-VLA-Hybrid-Architectures Β§2, Β§5).
  • It proves the fusion is interference-free at scale. BAGEL beats strong specialist models on both understanding and generation simultaneously β€” evidence that separate-QKV/FFN-shared-attention is a real solution to the multi-objective conflict, not a compromise.
  • Its emergent world-modeling is the bridge to WAMs. Future-frame prediction and world navigation emerging from pure multimodal pretraining is exactly the prior a WAM wants β€” so a BAGEL-initialized policy starts with a usable world model.

3. Architecture

BAGEL's Mixture-of-Transformer-Experts β€” an Understanding expert (Next-Token Prediction) and a Generation expert (Velocity Prediction), each with its own QKV + FFN but sharing one "Multi-modal Self-Attention"; separate Text Tokenizer + Und Encoder (semantic) and Gen Encoder (pixel/VAE) feed the two experts (architecture figure from arXiv 2505.14683, Β© the authors)

flowchart LR
  subgraph ENC[Dual visual encoders]
    VAE[VAE Β· pixel-level] 
    VIT[ViT Β· semantic-level]
  end
  IN[text Β· image Β· video Β· web] --> VAE
  IN --> VIT
  VAE --> G[Generation expert<br/>separate QKV/FFN]
  VIT --> U[Understanding expert<br/>separate QKV/FFN]
  U <-. shared attention layers .-> G
  U --> OU[text / understanding]
  G --> OG[image / edit / future frame]
Loading
  • MoT-Experts: understanding + generation, separate QKV + FFN per expert, shared attention. Init from Qwen2.5-7B-Instruct; 7B active / 14B total.
  • Dual vision path: VAE (pixel detail β†’ generation) + ViT (semantics β†’ understanding).
  • Objective: a "Next Group of Token Prediction" paradigm over interleaved multimodal tokens.
  • Data: trillions of interleaved text/image/video/web tokens across pretrain β†’ continued training β†’ SFT.

4. Results (paper-reported)

Task BAGEL Competitor
MME (understanding) 2388 Qwen2.5-VL 2347
MMBench 85.0 Qwen2.5-VL 83.5
GenEval (text→image) 0.88 SD3-Medium 0.74
GEdit-Bench (SC, editing) 7.36 Step1X-Edit 7.09
  • Emergent staging: understanding + generation appear early; basic editing next; complex/intelligent editing later β€” and free-form manipulation, multiview synthesis, and world navigation ("world-modeling" tasks) emerge with scale.

5. Significance & limitations

Significance. BAGEL is the load-bearing base of the three-expert hybrid family β€” the demonstration that a shared-attention, per-expert-QKV/FFN MoT with VAE+ViT dual encoders can unify understanding and generation at frontier quality. Robot VLAs (BagelVLA, HALO) inherit exactly this and add an action expert.

Limitations (from a VLA standpoint).

  1. Not a robot model. No action head, no embodiment β€” its world-modeling is emergent and qualitative, not benchmarked for control.
  2. Heavy (14B total). The base ticket before an action tower is even added.
  3. World-navigation / future-frame are demonstrations, not controllable dynamics β€” a policy still needs an action expert + robot data (what BagelVLA/HALO supply).
  4. General-purpose objective β‰  manipulation prior β€” the gap between "can imagine plausible frames" and "predicts contact-accurate dynamics" is the same one all pixel WAMs face (Review-World-Models Β§6).

6. Links

← Back to Reviews Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally