Skip to content

RSS 2026 HoMMI

hwoo.han edited this page Jul 25, 2026 · 2 revisions

HoMMI β€” Learning Whole-Body Mobile Manipulation from Human Demonstrations

Venue: RSS 2026 (Imitation Learning session) Β· Authors: Xiaomeng Xu, Jisang Park, Han Zhang, Eric Cousineau, Aditya Bhat, Jose Barreiros, Dian Wang, Jeannette Bohg, Shuran Song β€” Stanford Γ— Toyota Research Institute Category: Robot-free data collection + whole-body mobile manipulation Trend tag: RSS 2026 thread 2 β€” human data & cross-embodiment transfer arXiv: 2603.03243 Β· project Β· code

Compiled from the verified RSS 2026 abstract and the paper's Fig. 1.

Key figure

HoMMI overview (Figure 1 of arXiv 2603.03243, Β© the authors)

Figure 1 of the paper. (a) Scalable data collection: a human with dual UMI-style handheld grippers plus head-mounted egocentric sensing performs household tasks (towel-to-bin, cart pushing, cushion handling) β€” no robot in the loop. (b) The challenge: side-by-side human demonstration vs robot deployment showing the two gaps the policy design must bridge β€” the visual gap (human hands/grippers visible in egocentric views vs robot arms) and the kinematic gap (human head height/reach vs the shorter mobile bimanual robot). (c) Resulting skills: whole-body manipulation with active perception, long-horizon navigation while carrying, and bimanual coordination β€” all learned from the robot-free demonstrations.

Problem

UMI-style handheld interfaces made robot-free data collection work for tabletop arms β€” but mobile manipulation needs global context (where the body is going, what the head sees) that a wrist-mounted gripper camera can't provide, and adding egocentric sensing widens the human-to-robot embodiment gap in both observation and action spaces.

Method

  • Interface: UMI augmented with egocentric sensing β€” portable, robot-free, scalable capture of whole-body mobile-manipulation demonstrations.
  • Cross-embodiment hand-eye policy design to bridge the widened gap: an embodiment-agnostic visual representation, a relaxed head-action representation, and a whole-body controller that realizes commanded hand-eye trajectories through coordinated whole-body motion under robot-specific physical constraints.
  • Full code, data, and hardware design release stated.

Results (as reported)

  • Enables long-horizon mobile manipulation requiring bimanual and whole-body coordination, navigation, and active perception β€” trained entirely from robot-free human demonstrations.

Significance

The Song/Bohg/TRI lineage (UMI β†’ this) extending the robot-free data thesis from arms to whole bodies β€” the practical complement to Ξ¨β‚€'s teleop-based humanoid pipeline and EgoHumanoid's (#204) egocentric loco-manipulation co-training. The design insight worth retaining: the fix for a bigger embodiment gap is not more data but interface-level abstraction (hand-eye trajectories + a constraint-aware whole-body controller as the translator) β€” the mobile-manipulation echo of the canonical-representation moves seen across RSS 2026. See Review-Humanoid-VLA and CoRL-2025-DexUMI for the interface lineage.

← RSS 2026 survey Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally