-
Notifications
You must be signed in to change notification settings - Fork 0
NeurIPS 2025 VideoVLA
Venue: NeurIPS 2025 Β· arXiv: 2512.06963 Β· Project: https://videovla-nips2025.github.io/ Category: World Models / VLA Architecture
flowchart LR
Obs[Obs + language] --> DiT[Multimodal DiT]
DiT --> Action[Action chunk]
DiT --> Video[Future video]
Action & Video --> Loss[Joint training:<br/>action + video prediction]
Video -.correlation.-> Success[Imagined futures that<br/>resemble success β real success]
A VLA that predicts only actions throws away a huge signal: what the world should look like after successful execution. Prior video-prediction + policy work (e.g., DreamGen) trains them separately or uses video as an auxiliary task. VideoVLA asks: can we use a pretrained video generator directly as the VLA backbone?
Multimodal DiT (Diffusion Transformer), built on the pretrained CogVideoX-5B video generator adapted into a Video-Action Diffusion Transformer by adding actions as a new output modality. It jointly predicts:
- Action chunk (tokenized continuous actions)
- Future video (frames conditioned on the chosen action)
Both are generated in the same diffusion process (DDPM loss with synchronous noise scheduling on video latents and action tokens), conditioned on current observation + language. Crucially, the video backbone is pretrained on web-scale video before the action head is added β the claim is that video generation priors transfer to manipulation. There is no separate VLM; the video generator itself supplies the vision-language priors.
Evaluated on SIMPLER (Google Robot, WidowX) and a real-world Realman 7-DoF arm, against OpenVLA, RT-1/2-X, Octo, SpatialVLA, Οβ, and CogACT.
- SIMPLER Google Robot (Visual Matching): 80.4% average β vs 12.6% when training CogVideoX from scratch, underscoring the importance of the pretrained video prior.
- Generalizes to new embodiments and novel objects: SIMPLER novel-object 65.2%, cross-embodiment transfer 48.6%; real-robot in-domain 64.6%, novel-object 50.6%, cross-embodiment 58.0%.
- Imagined futures have strong correlation with task success β a free interpretability signal (human eval: imagined-future success 84.0% vs 65.2% actual execution on novel objects).
- Competitive with specialized VLAs on manipulation benchmarks despite no VLM in the pipeline.
A Video-Action Model (VAM) β VideoVLA is the NeurIPS 2025 version of the "video backbone replaces the VLM" thesis that ICLR 2026's mimic-video and related work push further (see Review: VLA Architectures Β§5.E).
Together with SAMPO (scale-wise AR) and OSVI-WM (one-shot visual imitation via world model), VideoVLA defines the NeurIPS 2025 world-model-first design direction.
Lineage: DreamGen (CoRL 2025) β VideoVLA + DreamVLA (NeurIPS 2025) β Cosmos Policy / Ctrl-World (ICLR 2026).
- arXiv: https://arxiv.org/abs/2512.06963
- NeurIPS virtual page: https://neurips.cc/virtual/2025/loc/san-diego/poster/117768
- Project: https://videovla-nips2025.github.io/
- DreamGen (CoRL 2025 ancestor)
- DreamVLA (world-knowledge sibling)
- Cosmos Policy Β· Ctrl-World (ICLR 2026 descendants)
- Review: VLA Architectures β Β§5.E
β Back to NeurIPS-2025
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)