Skip to content

CVPR 2026 DM0

Heungwoo edited this page Jun 1, 2026 · 2 revisions

DM0 β€” An Embodied-Native VLA towards Physical AI

Venue: CVPR 2026 Category: VLA Architecture (Embodied-native pretraining) Trend tag: Trend 1 Team: Dexmal & StepFun (project leads Erjin Zhou, Tiancai Wang) Β· arXiv Feb 16 2026 Backbone: Qwen3-1.7B LLM + perception encoder; ~2B total params with a flow-matching action expert

Approach diagram

flowchart TB
  subgraph Pre["Stage 1: Pretraining (1.2T tokens)"]
    WEB["web text + image"]
    DRIVE["driving data"]
    EMBO["embodied logs"]
  end
  Pre --> BB["DM0 VLM backbone<br/>(Qwen3-1.7B + perception encoder)"]
  BB --> MID["Stage 2: Mid-Training<br/>add flow-matching action expert<br/>(text + discrete + continuous actions)"]
  MID --> POST["Stage 3: Post-Training<br/>specialize per embodiment"]
  POST --> ESS["Embodied Spatial Scaffolding<br/>subtask β†’ goal bbox β†’ EE traj β†’ action tokens"]
  ESS --> ACT["action prediction"]
Loading

Problem

Most VLAs are language-pretrained then fine-tuned on robot data. This means the backbone learned its world prior from text and image data that has no embodied structure β€” no notion of contact, force, kinematics, or spatial scaffolding. The resulting models are good at language, weak at embodied reasoning.

Method

DM0 uses a three-stage pipeline on a Qwen3-1.7B-based VLM (with a perception encoder), ~2B params total:

  1. Pretraining (~1.2T tokens): unified large-scale training on heterogeneous data β€” (a) web text + image, (b) driving data (autonomous-driving trajectories with rich spatial structure), and (c) embodied logs β€” from the start, not as a fine-tune, so semantic knowledge and physical priors are acquired concurrently.
  2. Mid-Training: attaches a flow-matching action expert atop the VLM and jointly supervises text tokens, discrete action tokens, and continuous actions. A hybrid gradient strategy decouples the action-expert gradients from the VLM on embodied data (they are not backpropagated into the backbone) so the VLM keeps learning on non-embodied data and does not erode its general knowledge.
  3. Post-Training: specializes the model per target embodiment while retaining dialogue ability.

Embodied Spatial Scaffolding is a hierarchical supervision / structured information bottleneck applied in mid/post-training (not a pretraining objective). The model sequentially predicts subtask decomposition β†’ goal bounding boxes β†’ end-effector trajectory β†’ discrete action tokens, progressively constraining the action hypothesis space. DM0 unifies manipulation and navigation (navigation trajectories from Habitat are included).

Results

State-of-the-art on the RoboChallenge Table30 benchmark with only ~2B params:

Setting Model Success rate Task score
Specialist DM0 62.00% β€”
Specialist GigaBrain-0.1 51.67% β€”
Specialist Spirit-v1.5 51.00% β€”
Specialist Ο€0.5 42.67% β€”
Generalist DM0 37.3% 49.08
Generalist Ο€0.5 17.67% 31.27
Generalist Ο€0 9.0% 20.22

Significance

DM0 is the most explicit "embodiment-from-day-one" pretraining recipe published to date. Sits philosophically opposite to the standard "freeze a web VLM, bolt on an action head" pattern (GR00T series, Ο€0.7). If the trend holds, expect a 2026–2027 split between web-pretrained β†’ robot-fine-tuned (current dominant paradigm) and jointly embodiment-pretrained (DM0's bet).

Links

Related pages

← Back to CVPR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally