Skip to content

NeurIPS 2025 VideoVLA

Heungwoo edited this page Jun 1, 2026 · 2 revisions

VideoVLA β€” Video Generators Can Be Generalizable Robot Manipulators

Venue: NeurIPS 2025 Β· arXiv: 2512.06963 Β· Project: https://videovla-nips2025.github.io/ Category: World Models / VLA Architecture

Approach diagram

flowchart LR
  Obs[Obs + language] --> DiT[Multimodal DiT]
  DiT --> Action[Action chunk]
  DiT --> Video[Future video]
  Action & Video --> Loss[Joint training:<br/>action + video prediction]
  Video -.correlation.-> Success[Imagined futures that<br/>resemble success β†’ real success]
Loading

Problem

A VLA that predicts only actions throws away a huge signal: what the world should look like after successful execution. Prior video-prediction + policy work (e.g., DreamGen) trains them separately or uses video as an auxiliary task. VideoVLA asks: can we use a pretrained video generator directly as the VLA backbone?

Method

Multimodal DiT (Diffusion Transformer), built on the pretrained CogVideoX-5B video generator adapted into a Video-Action Diffusion Transformer by adding actions as a new output modality. It jointly predicts:

  • Action chunk (tokenized continuous actions)
  • Future video (frames conditioned on the chosen action)

Both are generated in the same diffusion process (DDPM loss with synchronous noise scheduling on video latents and action tokens), conditioned on current observation + language. Crucially, the video backbone is pretrained on web-scale video before the action head is added β€” the claim is that video generation priors transfer to manipulation. There is no separate VLM; the video generator itself supplies the vision-language priors.

Results

Evaluated on SIMPLER (Google Robot, WidowX) and a real-world Realman 7-DoF arm, against OpenVLA, RT-1/2-X, Octo, SpatialVLA, Ο€β‚€, and CogACT.

  • SIMPLER Google Robot (Visual Matching): 80.4% average β€” vs 12.6% when training CogVideoX from scratch, underscoring the importance of the pretrained video prior.
  • Generalizes to new embodiments and novel objects: SIMPLER novel-object 65.2%, cross-embodiment transfer 48.6%; real-robot in-domain 64.6%, novel-object 50.6%, cross-embodiment 58.0%.
  • Imagined futures have strong correlation with task success β€” a free interpretability signal (human eval: imagined-future success 84.0% vs 65.2% actual execution on novel objects).
  • Competitive with specialized VLAs on manipulation benchmarks despite no VLM in the pipeline.

Significance

A Video-Action Model (VAM) β€” VideoVLA is the NeurIPS 2025 version of the "video backbone replaces the VLM" thesis that ICLR 2026's mimic-video and related work push further (see Review: VLA Architectures Β§5.E).

Together with SAMPO (scale-wise AR) and OSVI-WM (one-shot visual imitation via world model), VideoVLA defines the NeurIPS 2025 world-model-first design direction.

Lineage: DreamGen (CoRL 2025) β†’ VideoVLA + DreamVLA (NeurIPS 2025) β†’ Cosmos Policy / Ctrl-World (ICLR 2026).

Links

Related pages

← Back to NeurIPS-2025

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally