-
Notifications
You must be signed in to change notification settings - Fork 0
Review MOTUS
Model: Motus β unified latent-action world model (WAM+VLA hybrid) Β· Tsinghua University (THU-ML) & collaborators (Jun Zhu / Hang Su group; incl. Horizon Robotics) Paper: arXiv 2512.13030 (v1 Dec 15 2025 Β· v2 Dec 25 2025) Β· open weights + code (GitHub thu-ml/Motus, HF motus-robotics/Motus, project page) Status: arXiv preprint (not yet peer-reviewed) β filed under Latest Papers. Authors: Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, β¦ Zhizhong Su, Lei Ma, Hang Su, Jun Zhu.
One of the concrete instances of the WAM+VLA hybrid NVIDIA's blog predicts (NVIDIA WAM thesis Β§2.6) β a single Mixture-of-Transformers that can be a world model, a VLA, an inverse-dynamics model, or a video generator on demand. Companion: VLA Architectures Β§4.2b Β· World Models Β· DYNA-2 Β· Being-H0.7.
- One model, five modes. Motus is a Mixture-of-Transformers (MoT) integrating three experts β understanding, video generation, action β with a UniDiffuser-style scheduler that flexibly switches among world model Β· VLA Β· inverse-dynamics model Β· video generation Β· video-action joint prediction. Unifying all of these in one pretraining is the contribution.
- Latent actions from optical flow. It learns latent actions by extracting a pixel-level "delta action" from optical flow, which lets it pretrain action representations at scale from action-free video (human + multi-robot).
- 8B params across four towers: video-generation model 5.00B, VLM 2.13B, action expert 641.5M, understanding expert 253.5M.
- SOTA on RoboTwin 2.0. On the RoboTwin 2.0 randomized multi-task setting: 88.66% vs X-VLA 72.80% and Ο0.5 42.98% (= +15% over X-VLA, +45% over Ο0.5). Real-world (two platforms, 9 tasks): +11β48% over Ο0.5.
- It operationalizes the "unified" corner of the WAM+VLA hybrid. Where DYNA-2 drops its video tower at inference (reactive) and Ο0.7 bolts a world model onto a VLA modularly, Motus keeps all modes live in one MoT and schedules which to run β the most literal reading of "collapse WAM and VLA into one pretraining."
- Optical-flow latent actions are a cheap, general action-supervision bridge. Like DYNA-2's hand-pose pseudo-actions, this turns action-free video into trainable action signal β but via a task-agnostic motion cue (flow) rather than hand tracking, so it extends to non-hand, multi-robot footage.
- Open weights + code. Unlike DYNA-2 (closed) or Cosmos 3 (partial), Motus is fully released, so its unified-MoT recipe is independently checkable.

Mixture-of-Transformers, four towers, one scheduler. Each modality is a specialized expert; the UniDiffuser-style scheduler sets per-modality noise levels so the same weights realize different models:
flowchart LR
IN[obs Β· language] --> UND[Understanding expert Β· 253M]
IN --> VGM[Video-generation expert Β· 5.0B]
IN --> ACT[Action expert Β· 641M]
VLM[VLM Β· 2.13B] --- UND
UND <-. shared self-attention .-> VGM
VGM <-. shared self-attention .-> ACT
UND <-. shared self-attention .-> ACT
SCHED[UniDiffuser scheduler<br/>sets per-modality noise level] -. selects mode .-> VGM
SCHED -. selects mode .-> ACT
ACT ==> OUT[action]
| Scheduler mode | What is conditioned on what | Equivalent model |
|---|---|---|
| clean obs β predict future frames | video from observation | world model |
| clean obs + language β action | action from obs+text | VLA |
| clean obs + clean future β action | action from a transition | inverse-dynamics model |
| noise β frames | unconditional video | video generator |
| joint denoise frames + action | co-generate both | video-action joint prediction |
Latent actions via optical flow. Motus derives a pixel-level "delta action" from optical flow between frames, giving an embodiment-agnostic latent action label learnable from any video β the substrate for large-scale action pretraining.
Three-phase training over a six-layer data pyramid.
- Phase 1 β Learning Visual Dynamics: adapt the video-generation model on multi-robot and human videos.
- Phase 2 β Learning Action Representations: pretrain the unified model with the optical-flow latent actions.
- Phase 3 β Specializing for the Target Robot: fine-tune on target-robot data.
- Data pyramid (quantity β, quality β): web data β egocentric human video β synthetic β task-agnostic β multi-robot trajectories β target-robot task data.

| Setting | Motus | X-VLA | Ο0.5 |
|---|---|---|---|
| RoboTwin 2.0 (randomized multi-task) | 88.66% | 72.80% | 42.98% |
- Real-world, two robot platforms, 9 tasks (fold towel Β· brew coffee Β· grind beans Β· pour water Β· touch keyboard Β· grab cube Β· place cube Β· get water from dispenser Β· put bread in oven): +11β48% over Ο0.5.
- Ablations report that unifying all functionalities/priors in one model is what drives the downstream gains (vs single-mode baselines).
Significance. Motus is the clearest open demonstration that a single scheduled MoT can serve every world/action role β the unified endpoint of Category E (Review-VLA-Architecture Β§4.2b). The optical-flow latent-action recipe is a general alternative to hand-pose (DYNA-2) or reconstruction-free latents (Ο-0).
Limitations.
- arXiv preprint, unreplicated. Numbers are the authors' own; no peer review yet.
- No explicit limitations/failure section. Future work only ("more universal motion priors; latent actions from internet-scale video").
- Optical flow as action proxy is coarse β flow conflates camera and object/effector motion; how well "delta action" grounds fine contact-rich control isn't isolated.
- RoboTwin 2.0 is simulation; real-world evidence is a 9-task, two-platform demo, not a broad hardware study.
- Paper: arXiv 2512.13030 Β· code/weights: GitHub Β· HF Β· project
- VLA Architectures Β§4.2b (WAMΓVLA hybrids) Β· World Models
- Sibling hybrids: DYNA-2 Β· Ο-0 Β· Being-H0.7 Β· Cortex 2.0 Β· Cosmos 3 / NVIDIA WAM
- Latest Papers
β Back to Latest Papers Β· Home Β· Reviews
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)