-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 Simulation Distillation
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: World Models & Memory Β· paper #17 Authors: Jacob Levy, Tyler Westenbroek, Kevin Huang, Fernando Palafox, Patrick Yin, Shayegan Omidshafiei, Dong-Ki Kim, Abhishek Gupta, David Fridovich-Keil arXiv: 2603.15759 Β· program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1 contrasts zero-shot sim-to-real failures (left column: peg insertion, table leg, slippery slope, foam) with successful executions after SimDist adaptation (right): the UR5e completes both precise assembly tasks and the Unitree Go2 traverses low-friction PTFE panels and memory foam, after only 15β30 minutes of real-world data.
End-to-end policy finetuning in new real-world environments is inefficient and brittle β model-free RL finetuning often collapses via catastrophic forgetting, especially on long-horizon contact-rich tasks. World models enable planning by counterfactual reasoning, but training action-conditioned robot world models directly in the real world requires diverse data at impractical scale. The question is how to obtain the coverage and supervision needed for planning-grade world models without collecting it in the real world.
SimDist pretrains a planning-oriented latent world model entirely in simulation and reduces real-world adaptation to supervised system identification. (1) A privileged state-based expert policy, its intermediate checkpoints, and its value function are trained with RL; diverse rollouts are generated by mixing expert and sub-optimal policies with contiguous action perturbations, giving dense reward/value supervision (100k trajectories for manipulation, 100M for the quadruped). (2) The world model β encoder, history encoder, transformer-based chunked latent dynamics predicting T future states in one forward pass, transformer sequence-to-sequence reward/value heads, base-policy head, and no pixel reconstruction β is pretrained on this data. (3) At deployment, MPPI planning runs with the encoder, reward model, and value function frozen; only the latent dynamics model is finetuned on real-world prediction losses, iterating collection and finetuning. Manipulation uses a UR5e with three 224Γ224 RGB cameras, ResNet-18 encoders, a 64-d latent, and H=T=5 at 5 Hz; the Go2 quadruped uses H=T=25 planned at 50 Hz on an RTX 4090M laptop.
Across four real-world tasks β Peg Insertion and Table Leg assembly (narrow/wide initial-condition grids), quadruped Slippery Slope (PTFE panels), and Foam β SimDist reliably improves with only 15β30 minutes of real-world data and typically reaches scores about 2Γ higher than any baseline (RLPD, IQL, SGFT-SAC, Diffusion Policy, Ο0.5), which stagnate or collapse during finetuning. Task throughput improves ~1.5β2Γ over zero-shot. On Slippery Slope the latent-dynamics loss on a held-out real trajectory drops from 0.076 (pretrained) to 0.019 (finetuned). Simulation ablations: full SimDist reaches 0.90/0.85 success (Peg/Table Leg) vs 0.10/0.05 with expert-only pretraining data; unfreezing the encoder or the value function during adaptation destroys performance, and adding pixel reconstruction drops manipulation to 0.32/0.21.
A crisp decomposition result for sim-to-real: task structure (representations, rewards, values) transfers across the dynamics gap, so only the dynamics model needs real-world correction β turning unstable real-world RL into stable supervised finetuning. Directly relevant to the planning-oriented vs generative world-model discussion in Review-World-Models.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)