-
Notifications
You must be signed in to change notification settings - Fork 0
CVPR 2026 EgoScale
Venue: CVPR 2026 Category: Egocentric data scaling / Dexterous Trend tag: Trend 5 Affiliations: NVIDIA (GEAR), UC Berkeley, University of Maryland
flowchart LR
EGO["20kh egocentric video"] --> LABEL["action-label via<br/>hand-pose extractor"]
LABEL --> EGOSCALE["EgoScale dataset"]
EGOSCALE --> MID["lightweight aligned<br/>human-robot mid-training"]
ROBOT["robot data"] --> MID
MID --> VLA["flow-based VLA<br/>(VLM backbone + DiT action expert,<br/>arch similar to GR00T N1)"]
VLA --> SCALE_LAW["log-linear scaling law<br/>(RΒ²=0.9983)<br/>ego-data volume vs. val loss"]
How much ego video do you actually need before robot performance improves? Prior work (a few hundred to a few thousand hours) was inconclusive on the scaling curve. EgoScale builds the dataset and runs the experiment.
- Curate 20,854 hours of action-labeled egocentric human video (>20Γ prior humanβrobot transfer efforts), labeled with wrist motion and retargeted dexterous hand actions.
- Pretrain a flow-based VLA policy (VLM backbone + DiT action expert, architecture similar to GR00T N1, with embodiment-conditioned MLP adapters) using a flow-matching objective.
- Two-stage transfer recipe: large-scale human pretraining β lightweight aligned human-robot mid-training β lightweight robot fine-tuning.
- Fit a log-linear scaling law (RΒ²=0.9983) relating ego-data volume to wrist/hand action validation loss, which in turn predicts real-robot success.
Near-perfect log-linear scaling (RΒ²=0.9983) of wrist/hand action validation loss with ego-video volume β a quantitative scaling law for egoβrobot pretraining at this scale, with validation loss predictive of real-robot success. The final policy improves average success rate by 54% over a no-pretraining baseline on a 22-DoF dexterous robotic hand, and transfers effectively to lower-DoF hands. Demonstrated long-horizon, one-shot-adaptable tasks include folding clothes, separating cards, and picking up fruit with tongs.
EgoScale's scaling law is the strongest empirical case that the next 10Γ of robot data is ego video. The ~20.9 kh dataset is the resource; the log-linear law (RΒ²=0.9983) is the proof; the 54% real-robot success-rate gain on a 22-DoF hand is the validation. Alongside UniDex and EgoVLA (which EgoScale cites as concurrent but smaller-scale human-data work), CVPR 2026 hosts the trio that establishes ego-video pretraining as a dominant 2026 data strategy.
- arXiv: 2602.16710
- Project: NVIDIA GEAR EgoScale page
- GR00T series review β VLA architecture is similar to GR00T N1 (flow-based, VLM backbone + DiT action expert)
- Cross-Embodiment review Β· Dexterous Manipulation review
- EgoDex Β· Human-Video Pretraining Β· UniDex Β· EgoVLA
- CVPR 2026 survey
β Back to CVPR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)