Skip to content

Review MOTUS

hwoo.han edited this page Aug 11, 2026 · 4 revisions

In-Depth Review β€” Motus: A Unified Latent Action World Model

Model: Motus β€” unified latent-action world model (WAM+VLA hybrid) Β· Tsinghua University (THU-ML) & collaborators (Jun Zhu / Hang Su group; incl. Horizon Robotics) Paper: arXiv 2512.13030 (v1 Dec 15 2025 Β· v2 Dec 25 2025) Β· open weights + code (GitHub thu-ml/Motus, HF motus-robotics/Motus, project page) Status: arXiv preprint (not yet peer-reviewed) β€” filed under Latest Papers. Authors: Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, … Zhizhong Su, Lei Ma, Hang Su, Jun Zhu.

One of the concrete instances of the WAM+VLA hybrid NVIDIA's blog predicts (NVIDIA WAM thesis Β§2.6) β€” a single Mixture-of-Transformers that can be a world model, a VLA, an inverse-dynamics model, or a video generator on demand. Companion: VLA Architectures Β§4.2b Β· World Models Β· DYNA-2 Β· Being-H0.7.


1. TL;DR

  1. One model, five modes. Motus is a Mixture-of-Transformers (MoT) integrating three experts β€” understanding, video generation, action β€” with a UniDiffuser-style scheduler that flexibly switches among world model Β· VLA Β· inverse-dynamics model Β· video generation Β· video-action joint prediction. Unifying all of these in one pretraining is the contribution.
  2. Latent actions from optical flow. It learns latent actions by extracting a pixel-level "delta action" from optical flow, which lets it pretrain action representations at scale from action-free video (human + multi-robot).
  3. 8B params across four towers: video-generation model 5.00B, VLM 2.13B, action expert 641.5M, understanding expert 253.5M.
  4. SOTA on RoboTwin 2.0. On the RoboTwin 2.0 randomized multi-task setting: 88.66% vs X-VLA 72.80% and Ο€0.5 42.98% (= +15% over X-VLA, +45% over Ο€0.5). Real-world (two platforms, 9 tasks): +11–48% over Ο€0.5.

2. Why it matters

  • It operationalizes the "unified" corner of the WAM+VLA hybrid. Where DYNA-2 drops its video tower at inference (reactive) and Ο€0.7 bolts a world model onto a VLA modularly, Motus keeps all modes live in one MoT and schedules which to run β€” the most literal reading of "collapse WAM and VLA into one pretraining."
  • Optical-flow latent actions are a cheap, general action-supervision bridge. Like DYNA-2's hand-pose pseudo-actions, this turns action-free video into trainable action signal β€” but via a task-agnostic motion cue (flow) rather than hand tracking, so it extends to non-hand, multi-robot footage.
  • Open weights + code. Unlike DYNA-2 (closed) or Cosmos 3 (partial), Motus is fully released, so its unified-MoT recipe is independently checkable.

3. Architecture

Motus architecture β€” three experts (Video Gen. Model Β· Action Expert Β· Understanding Expert) coupled by a shared "Tri-modal Joint Attention", each with its own AdaLN/LayerNorm + QKV + FFN; video/action encoders & decoders on the outside, a frozen pre-trained VLM feeding the understanding expert (architecture figure from arXiv 2512.13030, Β© the authors)

Mixture-of-Transformers, four towers, one scheduler. Each modality is a specialized expert; the UniDiffuser-style scheduler sets per-modality noise levels so the same weights realize different models:

flowchart LR
  IN[obs Β· language] --> UND[Understanding expert Β· 253M]
  IN --> VGM[Video-generation expert Β· 5.0B]
  IN --> ACT[Action expert Β· 641M]
  VLM[VLM Β· 2.13B] --- UND
  UND <-. shared self-attention .-> VGM
  VGM <-. shared self-attention .-> ACT
  UND <-. shared self-attention .-> ACT
  SCHED[UniDiffuser scheduler<br/>sets per-modality noise level] -. selects mode .-> VGM
  SCHED -. selects mode .-> ACT
  ACT ==> OUT[action]
Loading
Scheduler mode What is conditioned on what Equivalent model
clean obs β†’ predict future frames video from observation world model
clean obs + language β†’ action action from obs+text VLA
clean obs + clean future β†’ action action from a transition inverse-dynamics model
noise β†’ frames unconditional video video generator
joint denoise frames + action co-generate both video-action joint prediction

Latent actions via optical flow. Motus derives a pixel-level "delta action" from optical flow between frames, giving an embodiment-agnostic latent action label learnable from any video β€” the substrate for large-scale action pretraining.

Three-phase training over a six-layer data pyramid.

  • Phase 1 β€” Learning Visual Dynamics: adapt the video-generation model on multi-robot and human videos.
  • Phase 2 β€” Learning Action Representations: pretrain the unified model with the optical-flow latent actions.
  • Phase 3 β€” Specializing for the Target Robot: fine-tune on target-robot data.
  • Data pyramid (quantity ↓, quality ↑): web data β†’ egocentric human video β†’ synthetic β†’ task-agnostic β†’ multi-robot trajectories β†’ target-robot task data.

Motus's six-layer data pyramid β€” from a broad, abundant base (Web Data β†’ Egocentric Human Videos β†’ Synthetic β†’ Task-Agnostic) up to scarce, high-quality tips (Multi-Robot β†’ Target-Robot Task Trajectory Data) (data-pyramid figure from arXiv 2512.13030, Β© the authors)


4. Results (paper-reported)

Setting Motus X-VLA Ο€0.5
RoboTwin 2.0 (randomized multi-task) 88.66% 72.80% 42.98%
  • Real-world, two robot platforms, 9 tasks (fold towel Β· brew coffee Β· grind beans Β· pour water Β· touch keyboard Β· grab cube Β· place cube Β· get water from dispenser Β· put bread in oven): +11–48% over Ο€0.5.
  • Ablations report that unifying all functionalities/priors in one model is what drives the downstream gains (vs single-mode baselines).

5. Significance & limitations

Significance. Motus is the clearest open demonstration that a single scheduled MoT can serve every world/action role β€” the unified endpoint of Category E (Review-VLA-Architecture Β§4.2b). The optical-flow latent-action recipe is a general alternative to hand-pose (DYNA-2) or reconstruction-free latents (Ο‰-0).

Limitations.

  1. arXiv preprint, unreplicated. Numbers are the authors' own; no peer review yet.
  2. No explicit limitations/failure section. Future work only ("more universal motion priors; latent actions from internet-scale video").
  3. Optical flow as action proxy is coarse β€” flow conflates camera and object/effector motion; how well "delta action" grounds fine contact-rich control isn't isolated.
  4. RoboTwin 2.0 is simulation; real-world evidence is a 9-task, two-platform demo, not a broad hardware study.

6. Links

← Back to Latest Papers Β· Home Β· Reviews

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally