Skip to content

CVPR 2026 EgoVLA

Heungwoo edited this page Jun 1, 2026 · 2 revisions

EgoVLA β€” Learning VLA Models from Egocentric Human Videos

Venue: CVPR 2026 Category: Egocentric VLA pretraining Trend tag: Trend 5 Affiliations: UCSD + UIUC + MIT + NVIDIA

Approach diagram

flowchart LR
  EGO["egocentric human video"] --> POSE["wrist pose + MANO<br/>hand parameters"]
  POSE --> PRE["pretrain NVILA-2B VLA"]
  PRE --> FT["robot fine-tune"]
  FT --> RETARGET["IK + hand retargeting<br/>to humanoid"]
  RETARGET --> POL["deployable VLA"]
Loading

Problem

Human ego video is abundant; robot teleop is not. But the gap between "what a human's wrist did" and "what the robot's end-effector should do" is non-trivial: different kinematics, different gripper.

Method

  • Backbone is NVILA-2B (compact VLM), chosen for vision-language understanding at small size.
  • Pretrain on egocentric human video to predict future wrist pose + MANO hand parameters β€” a unified human action space β€” from images, language, and proprioception. Pretraining uses ~500K image-action pairs from HOI4D, HOT3D, HoloAssist, and TACO.
  • Fine-tune on robot demonstrations; at deployment, human wrist+hand actions are mapped to the robot via inverse kinematics + hand retargeting rather than a single shared action space.

Results

Evaluated on the authors' Isaac Humanoid Manipulation Benchmark (NVIDIA Isaac Lab; Unitree H1 humanoid with two 12-DoF Inspire dexterous hands, 12 bimanual tasks: 7 short-horizon atomic + 5 long-horizon multi-stage). On seen backgrounds EgoVLA reaches 77.78% short-horizon vs 24.87% for the ACT baseline, and 45.93% long-horizon vs 2.22% for ACT. Human-video pretraining helps most on long-horizon and fine-grained tasks and on generalization to unseen backgrounds. Ablation: greater human-data diversity consistently improves generalization, but zero-shot deployment without robot fine-tuning yields 0% success.

Significance

EgoVLA is the cleanest demonstration that pretraining on ego video transfers to robot action prediction β€” provided the action spaces are made compatible. Sister paper to EgoScale (which scales the data) and UniDex (which scales across hands). Combined, the three define the 2026 ego-video manipulation-pretraining recipe.

Links

  • arXiv: 2507.12440
  • Project: rchalyang.github.io/EgoVLA

Related pages

← Back to CVPR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally