-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 mimic video
Venue: RSS 2026 (Imitation Learning session) Β· Authors: Jonas Pai*, Liam Achenbach*, β¦ Oier Mees, Elvis Nava β mimic robotics Γ ETH ZΓΌrich Γ Microsoft ZΓΌrich Γ UC Berkeley Β· arXiv: 2512.15692 Β· project Category: Video-action model (VAM) β alternative backbone class to VLAs Trend tag: RSS 2026 thread 3 β video/world models vs the VLA backbone
Compiled from the verified RSS 2026 abstract and the paper's Fig. 1; the video backbone is NVIDIA Cosmos-Predict2 (per the project page).

Figure 1 of the mimic-video paper. Top: the standard VLA pipeline β a VLM pretrained on static image-text pairs supplies semantics only, so large-scale robotics data must teach both dynamics and control in expensive post-training. Bottom: the Video-Action Model β a video model pretrained on video-text pairs already carries semantics plus visual dynamics, so small-scale robot data only has to teach control. Right: the payoff chart β success rate vs robot-data quantity, where the VAM at 10% of the data matches the VLA at 100% (the "10Γ sample efficiency" claim).
VLA backbones are pretrained on static, disconnected web data β semantically rich but "blind to physical causality." The policy must therefore infer dynamics and temporal dependencies entirely from robot trajectories, creating an unsustainable expert-data burden.
- Video-Action Model (VAM): pair a pretrained Internet-scale video model (which captures semantics and visual dynamics jointly) with a flow-matching action decoder conditioned on its latent representations.
- The decoder functions as an inverse dynamics model (IDM): the video model produces latent video-space action plans; the IDM translates them into low-level robot actions.
- Division of labor: pre-training handles physics + semantics; robot data only has to teach low-level control.
- State-of-the-art on simulated and real-world manipulation benchmarks.
- 10Γ sample efficiency and 2Γ convergence speed vs traditional VLA architectures.
The cleanest statement of the VAM thesis at RSS 2026 β that the right pre-training substrate for control is video, not vision-language. It converges with LDA-1B (dynamics in latent space), Cosmos-Policy (fine-tune a video model into a policy), and the WAM line (Review-WAM-vs-VLA-Robustness), and pressures the VLM4VLA question from the other side: if the vision encoder is the VLA bottleneck, a video-dynamics encoder may be the fix. The plan-in-video-latents + IDM decomposition is also the policy-side mirror of Qwen-RobotWorld's generation-side stack.
β RSS 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)