-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 Vid2World
Venue: ICLR 2026 Category: World Model / Video Diffusion Trend tag: Trend 8 (world models from internet-scale video) Affiliation: Tsinghua University Β· Chongqing University
flowchart LR
VD[Pretrained DynamiCrafter<br/>1.1B U-Net video diffusion] --> Caus["Causalization<br/>1) Causal mask on temporal attn<br/>2) Extrapolative weight transfer<br/> for temporal conv"]
Caus --> DF[Diffusion Forcing<br/>k_t ~ U 0,K independently per frame]
DF --> CAI[Causal Action Injection<br/>per-frame action embed via MLP]
CAI --> CAG[Causal Action Guidance<br/>action dropout p +<br/>classifier-free Ξ΅_guided = 1+Ξ» Ξ΅_cond β Ξ» Ξ΅_uncond]
CAG --> WM[Interactive autoregressive<br/>world model]
Two camps with complementary weaknesses:
- Domain-specific world models (DreamerV3, DIAMOND, NWM) predict accurately but require costly per-domain action-labeled data and produce low-fidelity rollouts.
- Pretrained video diffusion models (DynamiCrafter, Sora, Veo) generate high-fidelity video at internet scale but are non-interactive β they denoise full sequences with bidirectional temporal context, have no frame-level action conditioning, and cannot do autoregressive rollouts.
The paper argues the right move is "model-level transfer" of internet-scale video priors into a world model, rather than pretraining yet another world model on cross-domain action-labeled data (which is still data-hungry and yields low-fidelity output). Two technical barriers must be cleared: causal generation and frame-level action conditioning.
DynamiCrafter β a 1.1B U-Net-based video diffusion model pretrained on internet-scale videos (Xing et al., 2024). Vid2World post-trains it under a causal training objective and adds frame-level action conditioning.
Temporal attention layers. Causalize via causal masking (no parameter changes β attention is content-based, so restricting receptive field to past tokens is "free").
Temporal convolution layers. Symmetric kernels {w_t}_{t=-m}^{m} aggregate from past and future frames; the original kernels are non-causal. Three weight-transfer strategies are studied:
-
Shift Weight Transfer β shift the entire kernel m steps into the past, getting {w't}{t=-2m}^{0}. Preserves all weights but introduces temporal misalignment.
-
Masked Weight Transfer β keep only the {w_t}_{tβ€0} weights and zero the rest (hard causal mask at init). Causal but throws away future-facing weights.
-
Extrapolative Weight Transfer (proposed). Posits a linear feature relationship z_{t+k} β Ξ£ Ξ³_{k,j} z_{t-j} + Ξ²_k, then redistributes the future-side weights {w_i}_{i>0} onto the past side to preserve the original convolution output:
w'j = 1[jβ₯-m] Β· w_j + 1[-p+1β€jβ€0] Β· Ξ£ Ξ³{i,-j} w_i, b' = b + Ξ£ w_i Ξ²_i.
Detailed derivation in Appendix A.2; error bound (Proposition 1) in Appendix A.3.
Training objective for causal generation: Diffusion Forcing. Standard video diffusion uses a homogeneous noise level across frames (all frames share k). For causal autoregressive sampling, history frames must be clean (k=0) while the current frame is being denoised β a noise-level distribution the original training never sees. Vid2World adopts Diffusion Forcing (Chen et al., 2024): sample noise level independently per frame k_t ~ U(0,K). This exposes the model to all noise-level combinations and enables flexible causal rollouts.
Causal Action Injection. Frame-level action a_{t-1} is encoded via a lightweight MLP and added to the model's latent representation at temporal position t. This binds each predicted frame to its preceding action β the basis for fine-grained interactive control.
Action Dropout for Classifier-Free Guidance. Each timestep's action is independently dropped with probability p:
L(ΞΈ) = E [ Ξ£_t β Ξ΅_t β Ξ΅_ΞΈ([x^{k_Ο}Ο]{β€t}, [Γ£_Ο]{<t}, [k_Ο]{β€t}) βΒ² ], Γ£_t = β w.p. p, else a_t.
This gives both Ξ΅_cond = Ξ΅_ΞΈ(β¦, [a_Ο]{Ο<t}, β¦) (full action context) and Ξ΅_uncond = Ξ΅_ΞΈ(β¦, [a{Ο<t-1}, β ], β¦) (most recent action masked). Inference uses CFG-style:
Ξ΅_guided = (1+Ξ») Β· Ξ΅_cond β Ξ» Β· Ξ΅_uncond.
Theorem 4.1 (Causal Action Guidance as Probability Steering; proven in Appendix A.4): with H_t := ([x_Ο]{Ο<t}, [a_Ο]{Ο<t-1}) the history excluding the current action, this score composition is equivalent to sampling from a steered posterior pΜ(x_t | a_{t-1}, H_t) β p(x_t | H_t) Β· p(a_{t-1} | x_t, H_t)^Ο with Ο β (1+Ξ») β a history-consistent prior times an action-alignment classifier term. The guidance scale Ξ» β ββΊ is thus a knob trading off action responsiveness vs generation fidelity. (The appendix also contains Proposition 1 in A.3, the Extrapolative Weight Transfer error bound.)
On RT-1 robot manipulation: post-trained for 100k gradient steps, ~7 days on 4Γ A100 GPUs (extrapolative-weight-transfer variant). Two inference modes:
- Vid2World-NAR β denoise all frames simultaneously (matches baselines).
- Vid2World β autoregressive denoising with causal action guidance.
Ablation models in Table 2 are trained for only 30k gradient steps due to compute budget.
| Model | FVD β | FID β | SSIM β | LPIPS β | PSNR β | DreamSim β |
|---|---|---|---|---|---|---|
| Pre-trained Base Modelβ | 237.6 | 5.432 | 0.712 | 0.228 | 20.6 | β |
| Classifier Guidanceβ | 213.1 | 6.005 | 0.683 | 0.250 | 19.8 | 0.054 |
| ControlNetβ | 27.1 | 3.248 | 0.836 | 0.148 | 24.5 | β |
| Action-Conditionedβ | 24.2 | 2.965 | 0.852 | 0.134 | 25.6 | β |
| Language-Conditionedβ | 33.7 | 3.511 | 0.812 | 0.177 | 22.1 | β |
| AVIDβ | 39.3 | 3.436 | 0.842 | 0.142 | 25.3 | β |
| Vid2World-NARβ | 18.7 | 5.871 | 0.856 | 0.140 | 25.8 | 0.048 |
| Vid2World* | 18.5 | 5.806 | 0.842 | 0.152 | 24.6 | 0.054 |
β non-autoregressive prediction; *autoregressive prediction. Vid2World wins FVD in both modes, often by large margins; SSIM/LPIPS/PSNR competitive with or matching the best baselines.
3D Game Simulation β CS:GO (4 conditioning frames β autoregress to length 16; Pearce-Zhu 5.5M frames / 95h)
| Model | FVD β | FID β | SSIM β | LPIPS β |
|---|---|---|---|---|
| DIAMOND-Fast* | 577.1 | 115.6 | 0.449 | 0.547 |
| DIAMOND-HQ* | 368.5 | 87.2 | 0.447 | 0.510 |
| Vid2World* | 106.6 | 17.5 | 0.481 | 0.404 |
Relative gains over best DIAMOND configuration: β71.1% FVD, β79.9% FID.
| Model | FVD β | FID β | SSIM β | LPIPS β | PSNR β | DreamSim β |
|---|---|---|---|---|---|---|
| NWM (1B)β‘ single-step | 31.2 | 34.1 | 0.389 | 0.295 Β± 0.002 | 15.343 Β± 0.060 | 0.091 Β± 0.001 |
| NWM + Ego4D (1B)β‘ single-step | 41.0 | 34.9 | 0.361 | 0.368 Β± 0.003 | 14.072 Β± 0.075 | 0.138 Β± 0.002 |
| Vid2World* autoregressive | 59.4 | 42.9 | 0.481 | 0.3236 | 16.10 | 0.108 |
β‘single-step prediction (NWM is conditioned on the prediction timestep t directly, so it avoids autoregressive error accumulation). Vid2World is autoregressive over 16 frames + 4 history (context length 20 > training horizon of 16, demonstrating temporal generalization). On par with NWM and surpasses NWM+Ego4D on 4 of the 6 metrics (the paper's own claim), even under autoregressive error accumulation β it wins SSIM/LPIPS/PSNR/DreamSim (losing only FVD/FID) and posts the best SSIM in the block.
Vid2World rolls out three RT-1 checkpoints (Begin / 15% / Converged) inside the world model on the close-drawer task. Human evaluators annotate trajectory success. The success-rate ranking inside Vid2World matches the real-world ranking β a direct demonstration that the world model is good enough to use as a SIMPLER-style evaluation environment (Algorithm 3 in paper).
| WT variant | AG | FVD β | FID β | SSIM β | LPIPS β | PSNR β |
|---|---|---|---|---|---|---|
| Shift | β | 29.9 | 7.85 | 0.799 | 0.185 | 21.5 |
| Masked | β | 29.4 | 7.07 | 0.824 | 0.169 | 22.9 |
| Extrapolative | β | 28.6 | 7.52 | 0.832 | 0.162 | 23.4 |
| Masked | β | 25.8 | 6.84 | 0.840 | 0.159 | 23.9 |
| Extrapolative | β | 22.4 | 6.16 | 0.839 | 0.159 | 23.9 |
Both causalization (Masked > Shift, Extrapolative β Masked but slightly better) and action guidance (β rows beat β rows by 4β6 FVD) contribute independently.
PSNR / SSIM / LPIPS / DreamSim plotted as a function of Ξ» β [1.0, 4.0]. Performance improves with Ξ» initially (alignment to action helps) but degrades at high Ξ» due to over-sharpening. Sweet spot is around Ξ» β 2.
- Backbone scale. DynamiCrafter at 1.1B is "relatively lightweight." Authors expect larger video-diffusion backbones (NVIDIA Cosmos, Wan, etc.) to deliver better world-model fidelity but did not test due to compute.
- Training time. Causalization + action conditioning take 100k gradient steps Γ ~7 days Γ 4 A100. Authors hope future methods can do this in fewer steps.
- No explicit policy learning. Vid2World demonstrates Real2Sim evaluation, but the paper does not yet show that policies learned inside Vid2World transfer back to the real world (RL or model-based planning is left as future work).
- vs DIAMOND (Alonso et al. 2024): DIAMOND trains a domain-specific autoregressive diffusion world model for CS:GO. Vid2World, starting from a generic video diffusion checkpoint, beats DIAMOND-HQ by β71.1% FVD / β79.9% FID β a large gap that argues for model-level transfer of video priors over from-scratch domain training.
- vs NWM (Bar et al. 2025): NWM is a 1B-parameter dedicated navigation world model, conditioned directly on prediction timestep (so it skips error accumulation). Vid2World matches NWM's single-step performance from autoregressive rollouts, despite being a generic transfer.
- vs AVID (Rigter et al. 2024): AVID adapts a frozen video model with an action adapter; Vid2World's full causalization + action injection beats it across robot manipulation metrics.
- vs Genie 2, Cosmos, Wan as world models: these are larger generic video models that could potentially be Vid2World-ized; the paper explicitly flags this as future work.
- vs Genie Envisioner: GE trains a unified video+action stack from scratch on manipulation data. Vid2World takes the opposite path β preserve a pretrained video diffusion checkpoint, bolt on causality and action guidance. Different bets on where the training cost should be paid.
- vs Ctrl-World / WMPO: these use world models for RL / policy improvement; Vid2World provides the substrate their successors might use.
- First systematic study of video-diffusion-to-world-model transfer β the three weight-transfer schemes (with the Extrapolative variant's error bound, Proposition 1 / Appendix A.3) and the causal action guidance mechanism (formalized as probability steering in Theorem 4.1 / Appendix A.4) are the lasting methodological contributions.
- OpenReview: https://openreview.net/forum?id=pFyzqbUiF9
- Project page: https://knightnemo.github.io/vid2world/
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)