Skip to content

IROS 2026 DreamMimic

hwoo.han edited this page Sep 7, 2026 · 2 revisions

IROS 2026 β€” DreamMimic: Visuomotor Whole-Body Loco-Manipulation via World Model

Venue: IROS 2026 (Pittsburgh) Β· paper #279 Β· Shanghai Jiao Tong University Β· Tsinghua University (Yin, Lai) Β· code to be released. Paper: arXiv 2608.22278. Representative of: humanoid whole-body loco-manipulation Γ— world models β€” a WM used as distillation supervision + representation, not as an inference-time planner. Companions: Humanoid VLA Β· World Models Β· VLA Hybrid Architectures Β· IROS 2026 survey.

DreamMimic β€” privileged specialist teachers (trained in simulation with privileged obs) distill into a visual generalist student via world-model-assisted distillation: the world model supplies latent knowledge (proprio+visual prediction) as a representation + multi-step supervision, and Performance-Conditioned Guidance balances teacher guidance vs student exploration (pipeline figure from Yin & Lai, arXiv 2608.22278, Β© the authors)

1. Problem

Vision-based whole-body loco-manipulation on humanoids is hard: partial observability, contact-rich dynamics, and learning long-horizon behaviors from high-dimensional visual input. The goal: a vision-based humanoid controller (no privileged state at deployment) that stays stable over long horizons.

2. Method

DreamMimic distills privileged teacher policies into vision-based student controllers via world-model-assisted distillation:

  • RSSM repurposed β€” instead of a Dreamer-style RSSM for planning, it learns predictive latent dynamics that serve as (a) a representation space and (b) a multi-step supervision signal, and exposes compact predictive features to the student to reduce long-term drift.
  • Interaction-Aware Prediction Heads (IAPH) β€” predict contact and object state to ground the latent in agent–object dynamics.
  • Reward Prediction Head (RPH) β€” predicts instantaneous task reward to anchor the representation to behaviorally relevant signals.
  • Performance-Conditioned Guidance (PCG) β€” a reward-driven adaptive distillation schedule scoring teacher and student to balance guidance vs exploration (prevents premature teacher annealing and over-interference).

3. Results

  • On OMOMO and BEHAVE datasets: improved manipulation, robust vision-based control without privileged info, and consistent cross-simulator transfer across humanoid platforms.

4. Why it matters (WAM lens)

DreamMimic is a clean IROS 2026 datapoint for the "WAM as training-time scaffold, not inference policy" trend (survey Β§5.1): the world model is a distillation teacher's representation + multi-step supervisor, never run to plan at deployment (the student is a plain vision policy). This is the humanoid-control cousin of AtomVLA (WM as offline critic) β€” both use the WM to shape training, keeping the deployed policy reactive. It also advances Humanoid VLA's whole-body thread with a concrete recipe for stabilizing visual policy distillation under contact-rich dynamics.

Limitations (reviewer): teacher–student distillation needs a privileged teacher (sim); evaluated on mocap-derived OMOMO/BEHAVE + sim transfer (not extensive real-humanoid hardware in the abstract); RSSM latent fidelity bounds the supervision quality.

5. Links

← Back to IROS 2026 survey Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally