-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 Interactive World Simulator for Robot
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: World Models & Memory Β· paper #18 Authors: Yixuan Wang, Rhythm Syed, Fangyu Wu, Mengchao Zhang, Aykut Onol, Jose Barreiros, Hooshang Nayyeri, Tony Dear, Huan Zhang, Yunzhu Li arXiv: 2603.08546 Β· program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1 in three panels: (left) real-world ALOHA robot interaction data across six tasks (mug grasping, rope routing, rope collecting, T pushing, box packing, pile sweeping); (middle) the action-conditioned video model rolled out autoregressively, with a realism-vs-speed scatter showing "Ours (15 FPS)" above Dreamer4, Cosmos, UVA and DINO-WM in PSNR, and 10-minute stable rollouts (t = 0 to 6000); (right) the two applications β near-flat success curves across 100% world-sim to 100% real training-data mixtures, and a task-score correlation plot between world-sim and real evaluation.
Action-conditioned video prediction models ("world models") are promising for robot policy training and evaluation, but existing ones are either too slow for interactive use (heavy multi-step diffusion needing enterprise GPUs) or drift and accumulate errors over long-horizon rollouts, so they cannot serve as faithful surrogates for demonstration collection or reproducible policy evaluation.
The Interactive World Simulator is built from a moderate-sized robot interaction dataset in two stages: (1) an autoencoder with a CNN encoder and a consistency-model decoder (CTM-style training) maps 128Γ128 RGB frames to compact 2D latents; (2) with the autoencoder frozen, an action-conditioned latent dynamics model β also a consistency model, instantiated as 3D-conv blocks with FiLM modulation and spatiotemporal attention β is trained with next-frame supervision, injecting small noise into observation contexts for robustness. Inference is autoregressive with a shifting fixed-length context window, giving stable rollouts of over 10 minutes at 15 FPS on a single RTX 4090. Data: ~600 play episodes per real task (~6 hours of collection each) on an ALOHA bimanual robot, plus 10,000 scripted episodes for a MuJoCo T-pushing task; the mug-grasping model is only 176.02 MB, trained in ~6 h (stage 1) + ~12 h (stage 2) on one H200. Humans teleoperate inside the simulator (keyboard or kinematic device) to collect synthetic demonstrations, and policies can be rolled out inside it for evaluation.
Over 192-step action-conditioned rollouts aggregated across seven tasks, it beats Cosmos, UVA, Dreamer4 and DINO-WM on all metrics, e.g. PSNR 25.82 vs 20.81 (Dreamer4) and 17.79 (DINO-WM), and FVD 243.20 vs 799.34 (Cosmos) and 1747β2213 for the rest. Policies (DP, ACT, Ο0, Ο0.5) trained on 100-episode mixtures from 100% simulator data to 100% real data perform comparably: DP scores 87.9% with pure simulator data vs 90.3% with pure real data; ACT 76.2% vs 73.6%; Ο0.5 rises from 73.1% to 88.8% with more real data. Policy scores in the simulator correlate strongly with real-world scores across four tasks (r = 0.8455β0.9908).
Shows that lightweight consistency-model world models can replace both the physical robot for demonstration collection and much of real-world evaluation, with quantified sim-to-real ranking fidelity β a practical answer to the evaluation bottleneck highlighted by the Large Behavior Models analysis. Connects to the wiki's Review-World-Models thread and to evaluation-methodology discussions in RSS 2026 survey.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)