Skip to content

RSS 2026 mimic video

hwoo.han edited this page Jul 25, 2026 · 2 revisions

mimic-video β€” Video-Action Models for Generalizable Robot Control Beyond VLAs

Venue: RSS 2026 (Imitation Learning session) Β· Authors: Jonas Pai*, Liam Achenbach*, … Oier Mees, Elvis Nava β€” mimic robotics Γ— ETH ZΓΌrich Γ— Microsoft ZΓΌrich Γ— UC Berkeley Β· arXiv: 2512.15692 Β· project Category: Video-action model (VAM) β€” alternative backbone class to VLAs Trend tag: RSS 2026 thread 3 β€” video/world models vs the VLA backbone

Compiled from the verified RSS 2026 abstract and the paper's Fig. 1; the video backbone is NVIDIA Cosmos-Predict2 (per the project page).

Key figure

VLA vs VAM (Figure 1 of arXiv 2512.15692, Β© the authors)

Figure 1 of the mimic-video paper. Top: the standard VLA pipeline β€” a VLM pretrained on static image-text pairs supplies semantics only, so large-scale robotics data must teach both dynamics and control in expensive post-training. Bottom: the Video-Action Model β€” a video model pretrained on video-text pairs already carries semantics plus visual dynamics, so small-scale robot data only has to teach control. Right: the payoff chart β€” success rate vs robot-data quantity, where the VAM at 10% of the data matches the VLA at 100% (the "10Γ— sample efficiency" claim).

Problem

VLA backbones are pretrained on static, disconnected web data β€” semantically rich but "blind to physical causality." The policy must therefore infer dynamics and temporal dependencies entirely from robot trajectories, creating an unsustainable expert-data burden.

Method

  • Video-Action Model (VAM): pair a pretrained Internet-scale video model (which captures semantics and visual dynamics jointly) with a flow-matching action decoder conditioned on its latent representations.
  • The decoder functions as an inverse dynamics model (IDM): the video model produces latent video-space action plans; the IDM translates them into low-level robot actions.
  • Division of labor: pre-training handles physics + semantics; robot data only has to teach low-level control.

Results (as reported)

  • State-of-the-art on simulated and real-world manipulation benchmarks.
  • 10Γ— sample efficiency and 2Γ— convergence speed vs traditional VLA architectures.

Significance

The cleanest statement of the VAM thesis at RSS 2026 β€” that the right pre-training substrate for control is video, not vision-language. It converges with LDA-1B (dynamics in latent space), Cosmos-Policy (fine-tune a video model into a policy), and the WAM line (Review-WAM-vs-VLA-Robustness), and pressures the VLM4VLA question from the other side: if the vision encoder is the VLA bottleneck, a video-dynamics encoder may be the fix. The plan-in-video-latents + IDM decomposition is also the policy-side mirror of Qwen-RobotWorld's generation-side stack.

← RSS 2026 survey Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally