-
Notifications
You must be signed in to change notification settings - Fork 0
Review Omega0
In-Depth Review β Ο-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
Paper: Ο-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation Authors: Zhe Li*β , Zhenzhe Zhang*, Yangyang Wei*, Wenjie Zhang*, Xichen Yuan*, Peiyuan Zhi, Gen Li, Xinying Guo, Fengjie Gao, Jianfei Yangβ£, Shanghang Zhangβ£ (* equal, β project lead, β£ corresponding) Affiliations: MARS Lab (NTU) Β· Peking University Β· BAAI Β· HKUST(GZ) arXiv: 2608.06375 (v1 Aug 6, 2026, cs.RO) Β· Project: gentlefress.github.io/OMEGA-0_page Status: preprint (posted Aug 2026) β indexed here under Latest Papers
Companion reviews: Humanoid VLA Β· World Models Β· DreamZero Β· Ξ¨β Β· System 0/1/2 Β· Human Video β Robot Transfer Β· Real-Time Execution.
- A whole-body World Action Model for concurrent humanoid loco-manipulation. Most humanoid stacks decompose "move, then manipulate"; Ο-0 learns a single policy that steps, leans, balances, reaches, and manipulates simultaneously β from a language instruction, current multi-view observation, and proprioceptive state β outputting controller-compatible whole-body action latents executed by the SONIC low-level controller.
- The key design choice: future prediction as a reconstruction-free latent objective, not video generation. Unlike video-centered humanoid WAMs (MotionWAM, DiT4DiT) that make a predicted video trajectory the intermediate representation for action, Ο-0 predicts compact future observation embeddings (V-JEPA / Wan-latent targets) as a lightweight auxiliary signal, and directly denoises actions. Rationale: real humanoid vision is noisy/occluded/viewpoint-shifting during locomotion, so binding actions to a predicted video amplifies temporal inconsistencies into unstable whole-body motion β and better pixels don't imply better control.
- Prefix-guided dual-query attention couples foresight to action. A joint video-action latent predictor runs motion queries (one per action step) and video queries (future visual latents) with token-specific RoPE (2D for visual prefix, 3D for video queries, 1D for temporal action queries); motion queries attend to video queries so predicted scene-evolution cues are injected into the action representation β coordinated loco-manipulation without separate navigation and arm-control modules.
- Scales human/public motion into humanoid supervision via SONIC replay. A three-stage pipeline: (1) whole-body FAST-tokenized action VLM pretraining on Qwen3-VL-2B; (2) human-to-humanoid action-latent pretraining where public human motions (ARCTIC, Xperience-10M, Motion-X) are replayed in simulation by SONIC to produce robot-executable action latents + proprioceptive states (untrackable motions filtered out); (3) real-world fine-tuning with training-time Real-Time Chunking for smooth receding-horizon execution.
- Headline result: a single Ο-0 model does 11 real household tasks, ~2Γ the best baseline. On the Ο-HOME task suite, Ο-0 reaches 81.8% success / 90.3% task progress vs the strongest baseline Ο-0 (44.5% / 59.6%) and video-WAM DiT4DiT (43.6% / 61.0%) β all trained on the same real data, one multi-task policy each.
- It targets the axis humanoid VLAs punt on: concurrent whole-body coordination. Wiping a large table, mopping a floor, or loading a low washing-machine compartment fails if you separate locomotion, balance, and manipulation into phases. Ο-0 makes whole-body coordination the learned representation, extending the Humanoid VLA frontier past the "walk there, then stand and manipulate" decomposition of Ξ¨-0/AMO-style stacks.
- It is a latent WAM β the third design point vs DreamZero and MotionWAM. DreamZero jointly denoises pixel video + action in a 14B backbone; MotionWAM conditions on a video world model's denoising features. Ο-0 argues that for real-time humanoid control the policy "mainly needs compact future information," so it drops pixel-space video entirely from the action path (video decode is optional visualization only). This makes it the reconstruction-free / latent-predictive corner of the WAM design space in Review-World-Models.
- It operationalizes human-video transfer for whole-body humanoids β the human-video fork applied to legs+torso+arms, using controller-replay grounding (SONIC) to convert action-free human motion into robot-executable latents rather than co-training raw or synthesizing pixels.
- It ships a dataset the field lacks: Ο-HOME. 40.3 h / 4,827 episodes / 24 tasks at 30 Hz of household humanoid demonstrations with six synchronized modalities (egocentric RGB, exocentric RGB-D, whole-body SMPL motion, robot state, whole-body action latents, language) β closing a real gap for whole-body loco-manipulation training and evaluation.

Figure 2 of the paper. Stage 1 (left): a whole-body VLM (Qwen3-VL-2B) is fine-tuned to autoregressively predict FAST-tokenized whole-body action tokens from text + ego/exo visual tokens, giving an action-aware semantic prior. Stage 2 & 3 (center): a Joint Video-Action Latent Predictor fuses the VLM feature, T5 text, V-JEPA visual features, a view token, motion queries, and video queries; its future-aware motion feature conditions an Action DiT (0.45B) that denoises SONIC-compatible whole-body action latents, while video queries are supervised against frozen-Wan future latents. Right: the prefix-guided dual-query attention β self-attention on prefix/video/action queries, cross-attention of both query sets to the prefix, then motion queries attend to video queries β with 2D/3D/1D RoPE per token type. The predicted latents run on the SONIC whole-body controller in receding-horizon.
Given language β, a current view observation o^v_t, and robot state s_t, Ο-0 predicts a future chunk of whole-body action latents z_{t:t+H} executed by SONIC. State is compact: joint positions q_pos, dexterous-hand joints q_hand, and torso orientation as a continuous 6D rotation (quaternion converted to 6D to avoid the double-cover discontinuity); IMU linear acceleration / angular velocity are deliberately discarded.
| Stage | What trains | Objective | Data |
|---|---|---|---|
| 1 Β· Whole-body action VLM | Qwen3-VL-2B (fine-tuned) + whole-body FAST tokenizer | Next-token prediction of discrete whole-body action tokens; L1 tokenizer reconstruction | ARCTIC + Xperience-10M + Motion-X, unified to SMPL |
| 2 Β· Humanβhumanoid action-latent pretraining | joint predictor, state encoder, condition-fusion, Action DiT (VLM/V-JEPA/Wan frozen) | xβ-prediction action denoising + L_video (future latent MSE vs frozen Wan) |
Public motions replayed by SONIC β robot-executable latents + states; untrackable motions filtered |
| 3 Β· Real-world fine-tuning | same modules | Stage-2 objective + training-time RTC (random clean prefix M, loss on non-prefix only) | Ο-HOME real humanoid data |
Why the latent objective: future visual prediction is "intentionally lightweight and only serves as an auxiliary predictive objective" β the Action DiT denoises action latents directly from language/visual/state/future-aware conditions, with no test-time video-to-action inversion and no large video generator in the loop.
View tokens distinguish egocentric RGB, exocentric RGB, exocentric depth β exocentric RGB-D gives richer whole-body/scene supervision at training time while the robot deploys from first-person egocentric feedback. Deployment uses RTC: the last few frames of the previous chunk are cached as a clean prefix, so newly predicted chunks stay temporally consistent (the wiki's Real-Time Execution "trained-in continuation" pattern, here as a humanoid WAM).
- 40.3 hours / 4,827 episodes / 24 tasks @ 30 Hz, household scenarios across 8 capability groups (object retrieval, surface cleaning, appliance interaction, container transfer, cloth handling, storage arrangement, mobile manipulation, tool-based floor operation).
- Six synchronized modalities per episode: egocentric RGB (robot head cam = deployment view), exocentric RGB-D (ZED), proprioceptive state, whole-body SMPL motion references, whole-body action latents, language.
- Teleoperation rig: Pico 4 Ultra headset + handheld triggers + foot-mounted Pico trackers β head/hand/lower-body cues retargeted to a humanoid with Inspire DexHands; SONIC serves as the teleoperation policy.
- Leak-free protocol: the 11 downstream evaluation tasks are excluded from the Stage-2 pretraining pool; a single multi-task model is fine-tuned over all 11 tasks (no per-task policies).
Setup: 11 real household loco-manipulation tasks (Table 1), 10 trials each, single multi-task policy per method, all trained on the same real data. Metrics: success rate, subtask score (max 41), task progress.
| Method (category) | Success β | Score/41 β | Task Progress β |
|---|---|---|---|
| ACT (classical IL) | 8.2 | 10.6 | 32.4 |
| Diffusion Policy | 15.5 | 14.8 | 40.6 |
| Ο0.5 (VLA) | 27.3 | 20.9 | 52.8 |
| InternVLA-M1 | 31.8 | 21.8 | 55.6 |
| EgoVLA | 25.5 | 18.6 | 49.1 |
| GR00T-N1.7 | 22.7 | 19.7 | 49.8 |
| Ο-0 (humanoid, arm-centric + AMO) | 44.5 | 23.6 | 59.6 |
| Fast-WAM | 37.1 | 22.3 | 57.8 |
| DiT4DiT (video WAM) | 43.6 | 23.1 | 61.0 |
| Ο-0 (Ego) | 79.1 | 35.8 | 88.7 |
| Ο-0 (Omni) | 81.8 | 36.7 | 90.3 |
Reading: classical IL handles short chunks but not long-horizon whole-body coordination; VLAs gain from pretraining but their action interfaces aren't whole-body; humanoid/WAM baselines improve but are either arm-centric (Ο-0, Fast-WAM) or video-generation-centered without a controller-compatible whole-body latent interface (DiT4DiT). Ο-0's ego-only variant already ~1.8Γ the best baseline, and adding exocentric supervision (Omni) adds a few more points. The paper reports promising held-out object/scene and human-transfer generalization qualitatively.
- Sharpens the WAM taxonomy. Placed against DreamZero (pixel-video + action, 14B, arm/mobile) and MotionWAM/DiT4DiT (video-world-model-conditioned humanoid), Ο-0 is the latent-predictive, reconstruction-free, whole-body corner: future prediction is a cheap auxiliary signal, not the action pathway. It's a concrete counter-argument to "improving robotics = improving video generation" (DreamZero's thesis) for the humanoid real-time regime, where the authors argue compact future info beats pixel fidelity.
- Controller-grounded human data is the scalable lever. SONIC replay converts abundant action-free human/public motion into robot-executable whole-body latents β a different resolution of the human-video fork than emergence/decoupling/pixel-synthesis, specific to whole-body humanoids.
- A missing dataset filled. Ο-HOME's whole-body, multi-view, SMPL-annotated household corpus is directly reusable and addresses the "humanoid evaluation lags arms by a generation" gap noted in Review-Humanoid-VLA.
- Depends on the SONIC controller. Ο-0 predicts SONIC-compatible latents and uses SONIC for both teleoperation and low-level execution; portability to other whole-body controllers is untested, and the action interface is defined by SONIC's trackable-motion distribution (untrackable motions are filtered out of training).
- Single embodiment / lab-scale human transfer. Results are on one humanoid; human-transfer generalization is shown but at small scale, and cross-embodiment robustness is qualitative.
- Absolute long-horizon success is still moderate. 81.8% is a large relative win, but on a bespoke 11-task suite with a self-defined subtask/progress protocol; no shared external humanoid benchmark exists to place it against other labs.
- The reconstruction-free thesis is argued and supported on this setup, but there's no head-to-head ablation isolating "latent future target vs pixel-video target" at matched scale on the same tasks β the case against video-centered WAMs is made partly by baseline comparison rather than a controlled swap.
- IMU dynamics (accel/angular vel) are discarded for stability; whether that caps highly dynamic behaviors is unexplored.
- As a very recent preprint (Aug 2026), numbers are v1 and unreplicated; treat as early-signal.
- arXiv: https://arxiv.org/abs/2608.06375 Β· Project: https://gentlefress.github.io/OMEGA-0_page/
- Humanoid VLA β the frontier this extends to concurrent whole-body coordination
- World Models β the WAM taxonomy; Ο-0 is the latent-predictive corner
- DreamZero β the pixel-video WAM counterpoint (and its "improve video β improve policy" thesis)
- Ξ¨β β the staged-training humanoid baseline (Ο-0) it more-than-doubles
- Human Video β Robot Transfer β controller-replay grounding as a fourth transfer route
- Real-Time Execution β training-time RTC for smooth receding-horizon control
- Latest Papers β the preprint tracker this is filed under
β Back to Latest Papers Β· Home Β· Reviews
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)