-
Notifications
You must be signed in to change notification settings - Fork 0
CoRL 2026 HiPHI
Venue: CoRL 2026 (Austin, TX, Nov 9β12) Β· Noitom Robotics. Paper: arXiv 2608.16222. Representative of: the humanoid data substrate β a 617.5 h high-precision mocap corpus that keeps improving whole-body policies with scale. Companions: Humanoid VLA Β· CoRL 2026 survey.

Humanoid intelligence must learn over an extremely diverse space of whole-body motions and physically grounded interactions. Existing data sources force a trade-off: internet video is broad but lacks precise physical state, while lab motion-capture sets have accurate state but narrow behavioral coverage. HiPHI targets that gap with a large, high-fidelity corpus that systematically maximizes coverage of the human motion and interaction manifold.
HiPHI is a 617.5-hour optical motion-capture dataset captured at 90 Hz (~200.1 M frames) from 132 performers, with sub-millimeter marker tracking. It splits into 371.8 h of whole-body movement and 245.7 h of humanβobject interaction, the latter covering 40 real physical objects across 12 categories with synchronized object trajectories and mesh-level 3D geometry. Coverage is organized using the FrameNet linguistic framework β 214 Frame-LU labels across 22 frames β to structure the behavioral space. A companion benchmark suite scores motion-space diversity, interaction grounding, object consistency, and downstream physical-AI applications.
- Coverage. Broader and more uniform kinematic coverage than prior sets (1620 vs 1438 occupied cells; 1443 vs 1114 effective occupancy) with a longer tail (14.1% vs 10.7%).
- Quality. Low ground penetration (8 mm), 98.1% non-conflict frames, 95.7% near-surface grounding.
- Matched-budget tracking. In physics-based imitation (DeepMimic on Unitree G1), HiPHI gives the highest success rates and fastest convergence at both 3 h and 20 h budgets; cross-dataset MPJPE keeps dropping as training data scales from 3 h to 300 h.
- Real hardware. Policies transfer to a real Unitree G1, executing running, sitting, crawling, carrying, flipping, and pulling.
HiPHI argues that the bottleneck for whole-body humanoid policies is a data substrate β precise physical state plus systematically broad behavioral coverage β and shows the payoff is monotone with scale, so more of the same data keeps helping. The object-interaction split with meshes and trajectories makes it usable for contact-rich manipulation and loco-manipulation, not just locomotion.
Limitations (reviewer): optical mocap in a studio is expensive and constrains scene/lighting/object realism; tracking results center on the Unitree G1, so cross-embodiment generality is unverified; real-hardware behaviors are demonstrated qualitatively rather than with success-rate tables.
- arXiv 2608.16222
- Survey: CoRL 2026 Β· Related: Humanoid VLA
β Back to CoRL 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)