-
Notifications
You must be signed in to change notification settings - Fork 0
Review BAGEL
Model: BAGEL-7B-MoT β open unified multimodal understanding+generation model Β· ByteDance-Seed Paper: "Emerging Properties in Unified Multimodal Pretraining" β arXiv 2505.14683 Β· HF ByteDance-Seed/BAGEL-7B-MoT Β· blog Why it's in this wiki: BAGEL is not a robot model β but its Mixture-of-Transformers recipe is the template that BagelVLA, HALO, and the whole three-expert vision-tower family build on. This page documents the base recipe so the robot pages can point to one source.
Family: VLA Hybrid Architectures Β· BagelVLA Β· HALO Β· Motus.
- One model, understanding + generation, via Mixture-of-Transformer-Experts. BAGEL runs two experts β multimodal understanding and visual generation β with separate QKV projectors and FFNs per expert but shared attention layers. 7B active / 14B total, initialized from Qwen2.5-7B-Instruct.
- Dual visual encoders β the key idea robot VLAs copy. A VAE supplies pixel-level features (for generation) and a ViT supplies semantic features (for understanding), so one model both sees and draws without the two objectives fighting.
- Trillions of interleaved tokens. Pre-trained on interleaved text Β· image Β· video Β· web data; capability appears in stages (understanding+generation early β basic editing β complex editing).
- Emergent "world-modeling." Beyond image gen/edit/understand, it shows future-frame prediction, free-form manipulation, multiview synthesis, and world navigation β the very capabilities a WAM needs, which is why robot hybrids adopt its backbone.
- It is the base recipe of the three-expert MoT. HALO states it is "inspired by BAGEL's harmonization of multimodal tasks"; BagelVLA literally initializes from BAGEL and adds a third (action) expert. The pattern β per-expert QKV/FFN + shared self-attention, plus VAE+ViT dual vision encoders β is what lets robot models bolt on a video/action tower without wrecking language grounding (Review-VLA-Hybrid-Architectures Β§2, Β§5).
- It proves the fusion is interference-free at scale. BAGEL beats strong specialist models on both understanding and generation simultaneously β evidence that separate-QKV/FFN-shared-attention is a real solution to the multi-objective conflict, not a compromise.
- Its emergent world-modeling is the bridge to WAMs. Future-frame prediction and world navigation emerging from pure multimodal pretraining is exactly the prior a WAM wants β so a BAGEL-initialized policy starts with a usable world model.

flowchart LR
subgraph ENC[Dual visual encoders]
VAE[VAE Β· pixel-level]
VIT[ViT Β· semantic-level]
end
IN[text Β· image Β· video Β· web] --> VAE
IN --> VIT
VAE --> G[Generation expert<br/>separate QKV/FFN]
VIT --> U[Understanding expert<br/>separate QKV/FFN]
U <-. shared attention layers .-> G
U --> OU[text / understanding]
G --> OG[image / edit / future frame]
- MoT-Experts: understanding + generation, separate QKV + FFN per expert, shared attention. Init from Qwen2.5-7B-Instruct; 7B active / 14B total.
- Dual vision path: VAE (pixel detail β generation) + ViT (semantics β understanding).
- Objective: a "Next Group of Token Prediction" paradigm over interleaved multimodal tokens.
- Data: trillions of interleaved text/image/video/web tokens across pretrain β continued training β SFT.
| Task | BAGEL | Competitor |
|---|---|---|
| MME (understanding) | 2388 | Qwen2.5-VL 2347 |
| MMBench | 85.0 | Qwen2.5-VL 83.5 |
| GenEval (textβimage) | 0.88 | SD3-Medium 0.74 |
| GEdit-Bench (SC, editing) | 7.36 | Step1X-Edit 7.09 |
- Emergent staging: understanding + generation appear early; basic editing next; complex/intelligent editing later β and free-form manipulation, multiview synthesis, and world navigation ("world-modeling" tasks) emerge with scale.
Significance. BAGEL is the load-bearing base of the three-expert hybrid family β the demonstration that a shared-attention, per-expert-QKV/FFN MoT with VAE+ViT dual encoders can unify understanding and generation at frontier quality. Robot VLAs (BagelVLA, HALO) inherit exactly this and add an action expert.
Limitations (from a VLA standpoint).
- Not a robot model. No action head, no embodiment β its world-modeling is emergent and qualitative, not benchmarked for control.
- Heavy (14B total). The base ticket before an action tower is even added.
- World-navigation / future-frame are demonstrations, not controllable dynamics β a policy still needs an action expert + robot data (what BagelVLA/HALO supply).
- General-purpose objective β manipulation prior β the gap between "can imagine plausible frames" and "predicts contact-accurate dynamics" is the same one all pixel WAMs face (Review-World-Models Β§6).
- Paper: arXiv 2505.14683 Β· weights: HF Β· blog: seed.bytedance.com
- Robot descendants: BagelVLA (adds action expert) Β· HALO (three-expert VLA) Β· Motus
- Family & theory: VLA Hybrid Architectures Β· VLA Architectures Β§4.2b Β· World Models
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)