-
Notifications
You must be signed in to change notification settings - Fork 0
IROS 2026 DreamMimic
Venue: IROS 2026 (Pittsburgh) Β· paper #279 Β· Shanghai Jiao Tong University Β· Tsinghua University (Yin, Lai) Β· code to be released. Paper: arXiv 2608.22278. Representative of: humanoid whole-body loco-manipulation Γ world models β a WM used as distillation supervision + representation, not as an inference-time planner. Companions: Humanoid VLA Β· World Models Β· VLA Hybrid Architectures Β· IROS 2026 survey.

Vision-based whole-body loco-manipulation on humanoids is hard: partial observability, contact-rich dynamics, and learning long-horizon behaviors from high-dimensional visual input. The goal: a vision-based humanoid controller (no privileged state at deployment) that stays stable over long horizons.
DreamMimic distills privileged teacher policies into vision-based student controllers via world-model-assisted distillation:
- RSSM repurposed β instead of a Dreamer-style RSSM for planning, it learns predictive latent dynamics that serve as (a) a representation space and (b) a multi-step supervision signal, and exposes compact predictive features to the student to reduce long-term drift.
- Interaction-Aware Prediction Heads (IAPH) β predict contact and object state to ground the latent in agentβobject dynamics.
- Reward Prediction Head (RPH) β predicts instantaneous task reward to anchor the representation to behaviorally relevant signals.
- Performance-Conditioned Guidance (PCG) β a reward-driven adaptive distillation schedule scoring teacher and student to balance guidance vs exploration (prevents premature teacher annealing and over-interference).
- On OMOMO and BEHAVE datasets: improved manipulation, robust vision-based control without privileged info, and consistent cross-simulator transfer across humanoid platforms.
DreamMimic is a clean IROS 2026 datapoint for the "WAM as training-time scaffold, not inference policy" trend (survey Β§5.1): the world model is a distillation teacher's representation + multi-step supervisor, never run to plan at deployment (the student is a plain vision policy). This is the humanoid-control cousin of AtomVLA (WM as offline critic) β both use the WM to shape training, keeping the deployed policy reactive. It also advances Humanoid VLA's whole-body thread with a concrete recipe for stabilizing visual policy distillation under contact-rich dynamics.
Limitations (reviewer): teacherβstudent distillation needs a privileged teacher (sim); evaluated on mocap-derived OMOMO/BEHAVE + sim transfer (not extensive real-humanoid hardware in the abstract); RSSM latent fidelity bounds the supervision quality.
- Official program: IROS 2026 (paper #279) Β· survey: IROS 2026
- Related: Humanoid VLA Β· World Models Β· AtomVLA Β· 3D FlowMatch Actor
β Back to IROS 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)