Repository navigation
Review Egocentric Video Pretraining
Question: how do you turn label-free, first-person human video into a pre-training signal for robot / VLA policies? Egocentric video is abundant and rich in handβobject interaction, but it has no action labels and no robot embodiment β the whole game is extracting a learnable, transferable signal. Companion reviews: Human Video β Robot Transfer (the emergence/decoupling/synthesis fork) Β· Dexterous-Hand Data Pyramid (the L1/L2 data tiers) Β· World Models Β· Cross-Embodiment.
- Motivation. Robot/teleop data is scarce, expensive, and often unnatural; egocentric human video "already exists at effectively unbounded scale and carries exactly what a manipulation policy needs β how scenes evolve, how objects respond to contact, and how a hand interacts with them" (DYNA-2).
- The obstacle. Human video has no action labels and a human, not robot, embodiment. Prior naive use "mostly yielded visual features rather than transferable, fine-grained manipulation behavior" (Being-H0). Every method below is a different answer to "what supervised/self-supervised target do I extract from the pixels?"
Recover a proxy action per frame so the video can be treated like demonstration data.
- Hand-pose β wrist + grasp. DYNA-2 derives wrist poses β end-effector trajectories and a thumbβindex aperture β grasp signal as pseudo-actions; EgoScale pretrains on wrist motion + retargeted dexterous-hand actions.
- 3D hand/finger keypoints. EgoDex pairs video with 3D hand + finger tracks; Dexterous Point Policy (2606.10614) extracts 3D keypoints and trains an autoregressive transformer over them (no robot data).
- Optical-flow "delta action". Motus turns optical flow into a pixel-level delta action for embodiment-agnostic action pretraining.
- Part-level motion tokenization. Being-H0 pretrains via physical instruction tuning with part-level motion tokens + perspective alignment.
Instead of a hand-crafted proxy, learn a latent action space so future prediction is action-conditioned.
- DreamDojo learns a continuous-latent-action world model from 44,000 h of human video, distilled to real-time (10.9 FPS) β "the core bottleneck is the scarcity of action labels in the abundant, diverse human video."
- Being-H0.7 inserts latent queries with a posterior/prior branch; UniVLA learns task-centric latent action tokens.
Use next-frame (or next-latent) prediction as the pretraining loss, so the model absorbs dynamics.
- DYNA-2 β a World-Action Model: joint next-frame + next-action prediction on ~1M h human video; Motus β a video-generation expert co-trained with action; DreamDojo β a robot world model from human video. Latent-space variants avoid pixels (Ο-0, V-JEPA-style).
Reconstruct the human hand-object interaction, then map it onto a robot hand.
- DO AS I DO β reconstruct 4D hand-object dynamics from monocular video β dynamics-aware retargeting; DexImit β monocular human video β bimanual dexterity.
Pull non-action supervision the robot also needs.
- EgoTactile β recover full-hand grasp pressure from egocentric video (EgoPressureDiff), addressing the force signal that RGB video normally lacks.
| Dataset / corpus | Scale | Note |
|---|---|---|
| EgoDex | 829 h, 194 tasks | Vision Pro egocentric dex, paired 3D hand+finger |
| EgoVerse | global "around-the-world" | tackles fragmentation + embodiment-gap/scaling questions |
| EgoScale corpus | 20,854 h action-labeled | the scaling-law substrate (22-DoF hand) |
| Being-H0 | large-scale human-hand video (Being-H0.7: 200k h + 15k robot) | part-level motion tokens |
| DYNA-2 (vendor) | ~1M h human ego video | no robot data in pretraining |
| DreamDojo | 44,000 h | latent-action world model |
| UniDex | 10M frames from ego datasets | 8 dex hands, robot foundation suite |
| EgoVLA | egocentric human video | VLA learned from ego video |
(External anchors: Ego4D, EPIC-Kitchens β the raw web-scale base under most of these.)
- EgoScale established the headline result: dexterous-manipulation error improves log-linearly with human-video hours (RΒ²=0.9983), and pretraining yields +54% over a no-pretraining baseline on a 22-DoF hand.
- DYNA-2 claims a 1kβ1M h humanβrobot transfer law (fits RΒ²β0.88β0.93) with no plateau β the strongest (vendor-reported) version of the thesis.
- Consistent message: egocentric-video hours are a genuine scaling axis for manipulation, at least into the thousands-to-million-hour range.
Extracting a signal is half the problem; crossing the embodiment gap is the other half. Three camps (full treatment: Human Video β Robot Transfer):
- Emergence β transfer emerges once robot pretraining is diverse enough (Human2Robot Emergence, PI: ~2Γ on human-only generalization).
- Decoupling β separate videoβrepresentation from robotβcontrol (Ξ¨β: 800 h human + 30 h robot beats 10Γ corpora).
- Synthesis β turn video into robot data via retargeting/reconstruction (Β§2-D).
The dominant recipe is two-stage: pretrain on egocentric video, then fine-tune on a small robot teleop tip β often < 1 hour to tens of hours (see the data-pyramid verdict Β§3b). Egocentric video is the base, not the whole stack.
- Retargeting fidelity β human hand β robot hand is the load-bearing bridge; kinematic-only retargeting drops force (data pyramid Β§4).
- Missing force/tactile β RGB video has none; only partial recovery so far (EgoTactile).
- Action-label ambiguity β monocular depth/hand-object distance is ambiguous (DO AS I DO limitations); pseudo-actions are noisy.
- Yield β most raw clips don't survive faithful reconstruction (DO AS I DO: ~4% survival) β abundant β usable.
- Morphology & viewpoint β human 5-finger β robot hand; egocentric viewpoint β robot camera.
- Evaluation β no shared benchmark isolates "how much did the video pretraining actually contribute."
- Scaling / pretraining papers: EgoScale Β· Being-H0 Β· Being-H0.7 Β· DYNA-2 Β· DreamDojo Β· UniDex Β· EgoVLA Β· Motus
- Datasets: EgoDex Β· EgoVerse Β· EgoTactile
- Retargeting / reconstruction: DO AS I DO Β· DexImit Β· Dexterous Point Policy (2606.10614)
- Companion reviews: Human Video β Robot Transfer Β· Dexterous-Hand Data Pyramid Β· World Models Β· Humanoid VLA
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)