-
Notifications
You must be signed in to change notification settings - Fork 0
CoRL 2025 DreamGen
Venue: CoRL 2025 Β· Author: NVIDIA (GEAR Lab) Β· arXiv: 2505.12705 Full title: DreamGen: Unlocking Generalization in Robot Learning through Neural Trajectories Category: World Models for Policy Trend tag: World models for data generation
flowchart LR
D[Real robot data] --> WM[Fine-tuned video world model]
WM -- generated videos --> IDM[IDM / latent action model]
IDM -- pseudo-actions --> NT[Neural trajectories]
NT -- supervised IL --> POL[Visuomotor policy]
POL -- deploy --> REAL[Real robot]
Real-world data is expensive and sim-to-real has a gap. Video world models are now good enough (Sora, Cosmos, VEO-class) that we can imagine visually plausible rollouts β if we could train a policy entirely inside those rollouts, we'd scale beyond teleoperation.
DreamGen is a simple yet effective 4-stage offline data-generation pipeline β not online policy-training-in-imagination. The video world model is not action-conditioned; actions are recovered afterward:
- Fine-tune a state-of-the-art image-to-video world model (Cosmos, WAN 2.1, etc.) on target-robot trajectories.
- Generate synthetic robot videos from an initial frame + language instruction (familiar or novel tasks/environments).
- Recover pseudo-actions for each video using a latent-action model or an inverse-dynamics model (IDM), yielding neural trajectories (synthetic video + action pairs).
- Train a visuomotor policy via standard imitation learning on the neural trajectories. Demonstrated with Diffusion Policy, Οβ, and GR00T N1 (GR00T N1 used for the headline generalization runs).
Trained on teleoperation data from a single pick-and-place task in one environment, DreamGen unlocks zero-shot behavior and environment generalization β a GR1 humanoid performs 22 new verbs across 10 new environments.
| Setting | Baseline | + DreamGen |
|---|---|---|
| New behaviors, seen env (GR1) | ~0% | 43.2% |
| New behaviors, unseen env (GR1) | ~0% | 28.5% |
| 4 GR1 humanoid tasks (avg) | 37% | 46.4% |
| 3 Franka tasks (avg) | 23% | 37% |
| 2 SO-100 tasks (avg) | 21% | 45.5% |
Scaling neural trajectories shows a log-linear improvement in policy success (0 β 240k trajectories in RoboCasa). DreamGen Bench (evaluating Cosmos, WAN 2.1, Hunyuan, CogVideoX) finds video-generation quality positively correlates with downstream policy success.
DreamGen is the production-scale signal that video world models move into the training pipeline β as a synthetic-data engine rather than an online imagination simulator. It reframes the world model as a way to expand effective robot data (~10Γ over real teleoperation) and is integrated into NVIDIA's GR00T stack (released as GR00T-Dreams). Threads into ICLR 2026 via Ctrl-World, Cosmos Policy, WorldGym, and VLA-RFT (RL-in-world-model).
- arXiv: https://arxiv.org/abs/2505.12705
- NVIDIA CoRL 2025: https://www.nvidia.com/en-us/events/corl/
β Back to CoRL-2025
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)