-
Notifications
You must be signed in to change notification settings - Fork 0
Review Being H07
Model: Being-H0.7 β latent world-action model (unified and reactive WAM+VLA hybrid) Β· BeingBeyond (Zongqing Lu group) Paper: arXiv 2605.00078 (Apr 30 2026) Β· project page Β· GitHub BeingBeyond/Being-H Status: arXiv preprint (not yet peer-reviewed) β filed under Latest Papers. Authors: Hao Luo, Wanpeng Zhang, Yicheng Feng, Sipeng Zheng, β¦ Zongqing Lu. Built on the pretrained VLA Being-H0.5.
The blog's "most sophisticated hybrid" (NVIDIA WAM thesis Β§2.6): it fuses world-model imagination with VLA deployability by predicting a latent reasoning state rather than pixels β landing in both hybrid corners at once (unified architecture, reactive inference). Companion: VLA Architectures Β§4.2b Β· World Models Β· DYNA-2 Β· Ο-0 Β· MOTUS.
- Latent queries as an explicit reasoning interface. Being-H0.7 inserts a small set of learnable latent queries between the multimodal context and the action tokens. These slots attend to instruction + observation history + robot state and form a compact latent state before actions are generated β a "think, then act" bottleneck without generating any future frames.
- Dual-branch posterior/prior β the key trick. A future-informed posterior branch (training-only) replaces the queries with embeddings from future observations; a deployable prior branch infers the latent state from the current context alone. Aligning the two teaches the prior to encode future-aware, action-useful structure from the present β with hidden-state alignment + regularization to prevent latent collapse.
- Reactive at inference. At deployment it discards the posterior branch and performs no visual rollout β like DYNA-2/Ο-0, it keeps the world-model benefit while staying a fast VLA-style policy.
- Backbones: understanding = InternVL3.5, action = Qwen3, visual encoders = V-JEPA 2.1 (context-frame encoder kept trainable). Pretrained on large-scale egocentric human video + robot demonstrations (reported ~200k h human + ~15k h robot).
- It reconciles imagination and deployability without pixels. Rather than choose between an in-path video WAM (strong prior, slow) and a bare VLA (fast, weak prior), Being-H0.7 predicts a latent future-informed state and drops the predictive machinery at inference β the cleanest "reactive latent WAM" recipe alongside Ο-0.
-
Posterior/prior is a principled version of "co-training dropped at inference." Where DYNA-2 simply omits
z_tfrom the action head, Being-H0.7 explicitly trains a prior to match a future-informed posterior β a Play-LMP-style latent-plan structure that gives the reactive bet a training objective. - Egocentric-video pretraining at scale, continuing the human-video fork that DYNA-2 and EgoScale push.
flowchart LR
CTX[instruction Β· obs history Β· robot state] --> Q[Learnable latent queries<br/>compact reasoning state]
Q --> A[Action tokens Β· Qwen3]
subgraph TRAIN[training only]
FUT[future observations] --> POST[Posterior branch<br/>queries β future embeddings]
end
POST -. align latent reasoning space<br/>+ anti-collapse reg .- Q
Q -->|deployable prior: no rollout| A
- Understanding tower: InternVL3.5. Action tower: Qwen3. Visual encoders: V-JEPA 2.1 (both), with the context-frame encoder trainable.
- Prior (deployable): infers latent reasoning state from current context only.
- Posterior (training-only): substitutes the latent queries with embeddings from future observations; the two branches are jointly aligned in latent reasoning space, with hidden-state alignment + lightweight regularization to avoid latent collapse.
- Inference: posterior discarded, no visual rollout β reactive.
| Benchmark | Being-H0.7 |
|---|---|
| LIBERO | 99.2% |
| LIBERO-plus (zero-shot / fine-tuned) | 82.1% / 84.8% |
| RoboTwin 2.0 Hard | 89.6% |
| RoboCasa-50 | 62.1% |
| GR1 | 49.2% |
| CALVIN (ABCDβD / ABCβD) | 4.67 / 4.48 tasks |
- Outperforms Ο0.5, Fast-WAM, and Being-H0.5 across most benchmarks.
- Real world: leads on all five ability-oriented task suites across three platforms β PND Adam-U, Unitree G1, Franka FR3.
Significance. Being-H0.7 is the strongest academic case that a latent, reactive world-action model can match in-path WAMs on task quality while deploying like a VLA β and it gives the "reactive" bet (DYNA-2, Ο-0) a principled posterior/prior training objective rather than a design omission. It sits at the intersection of both hybrid axes in Review-VLA-Architecture Β§4.2b (unified architecture, reactive inference).
Limitations.
- arXiv preprint, unreplicated; numbers are the authors' own.
- Action-generation focus. The paper notes it currently focuses on action generation rather than text-generation tasks β the VLM's language breadth under this bottleneck isn't stressed.
- Latent-collapse risk is designed around, not eliminated β the regularization is a mitigation; robustness of the prior/posterior gap across domains isn't fully characterized.
- Exact pretraining hours are reported at the abstract level (~200k h human + ~15k h robot); the per-source breakdown is less detailed than the benchmark tables.
- Paper: arXiv 2605.00078 Β· project: research.beingbeyond.com/being-h07 Β· code: GitHub
- VLA Architectures Β§4.2b (WAMΓVLA hybrids) Β· World Models Β· Human Video β Robot Transfer
- Sibling hybrids: DYNA-2 Β· Ο-0 Β· MOTUS Β· Cortex 2.0 Β· Cosmos 3 / NVIDIA WAM
- Latest Papers
β Back to Latest Papers Β· Home Β· Reviews
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)