Skip to content

CVPR 2026 EgoScale

Heungwoo edited this page Jun 1, 2026 · 2 revisions

EgoScale β€” Scaling Dexterous Manipulation from Egocentric Human Data

Venue: CVPR 2026 Category: Egocentric data scaling / Dexterous Trend tag: Trend 5 Affiliations: NVIDIA (GEAR), UC Berkeley, University of Maryland

Approach diagram

flowchart LR
  EGO["20kh egocentric video"] --> LABEL["action-label via<br/>hand-pose extractor"]
  LABEL --> EGOSCALE["EgoScale dataset"]
  EGOSCALE --> MID["lightweight aligned<br/>human-robot mid-training"]
  ROBOT["robot data"] --> MID
  MID --> VLA["flow-based VLA<br/>(VLM backbone + DiT action expert,<br/>arch similar to GR00T N1)"]
  VLA --> SCALE_LAW["log-linear scaling law<br/>(RΒ²=0.9983)<br/>ego-data volume vs. val loss"]
Loading

Problem

How much ego video do you actually need before robot performance improves? Prior work (a few hundred to a few thousand hours) was inconclusive on the scaling curve. EgoScale builds the dataset and runs the experiment.

Method

  • Curate 20,854 hours of action-labeled egocentric human video (>20Γ— prior humanβ†’robot transfer efforts), labeled with wrist motion and retargeted dexterous hand actions.
  • Pretrain a flow-based VLA policy (VLM backbone + DiT action expert, architecture similar to GR00T N1, with embodiment-conditioned MLP adapters) using a flow-matching objective.
  • Two-stage transfer recipe: large-scale human pretraining β†’ lightweight aligned human-robot mid-training β†’ lightweight robot fine-tuning.
  • Fit a log-linear scaling law (RΒ²=0.9983) relating ego-data volume to wrist/hand action validation loss, which in turn predicts real-robot success.

Results

Near-perfect log-linear scaling (R²=0.9983) of wrist/hand action validation loss with ego-video volume — a quantitative scaling law for ego→robot pretraining at this scale, with validation loss predictive of real-robot success. The final policy improves average success rate by 54% over a no-pretraining baseline on a 22-DoF dexterous robotic hand, and transfers effectively to lower-DoF hands. Demonstrated long-horizon, one-shot-adaptable tasks include folding clothes, separating cards, and picking up fruit with tongs.

Significance

EgoScale's scaling law is the strongest empirical case that the next 10Γ— of robot data is ego video. The ~20.9 kh dataset is the resource; the log-linear law (RΒ²=0.9983) is the proof; the 54% real-robot success-rate gain on a 22-DoF hand is the validation. Alongside UniDex and EgoVLA (which EgoScale cites as concurrent but smaller-scale human-data work), CVPR 2026 hosts the trio that establishes ego-video pretraining as a dominant 2026 data strategy.

Links

  • arXiv: 2602.16710
  • Project: NVIDIA GEAR EgoScale page

Related pages

← Back to CVPR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally