-
Notifications
You must be signed in to change notification settings - Fork 0
Review DoAsIDo
Paper: "Do as I Do: Dexterous Manipulation Data from Everyday Human Videos" β arXiv 2606.19333 (Jun 17 2026) Authors: Bhawna Paliwal, Haritheja Etukuru, William Liang, Pieter Abbeel, Nur Muhammad Mahi Shafiullah, Jitendra Malik Β· UC Berkeley / NYU What it is: a device-free pipeline that turns ordinary monocular RGB human videos into robot-executable dexterous-hand trajectories β the cheapest L1βL4 path in the Dexterous-Hand Data Pyramid.
Companions: Dexterous-Hand Data Pyramid Β· AnyDexRT (the retargeting-bridge counterpart) Β· Dexterous Manipulation Β· Human Video β Robot Transfer.
- Everyday video β dexterous robot data, no gloves or teleop. Do As I Do reconstructs 4D hand-object dynamics from a single monocular RGB clip (internet, egocentric, or generated) and dynamics-aware retargets it into actions for a real dexterous hand.
- Reconstruction stack: HaWoR (hand tracking) + SAM 3D (object mesh/pose) + a guided-diffusion object tracker (fix shape at an anchor frame, track pose via flow matching at inference) + centroid/gravity alignment (GeoCalib).
- Dynamics-aware retargeting by sampling-based (MPPI-style) optimization in MuJoCo Warp, robustified for noisy references with warmup steps, random force perturbation, and a transition reward.
- Results: retargeting success 25% β 71% on reconstructed in-the-wild video and 72% β 81% on clean OakInk2 MoCap (vs annealed-sampling alone); object tracking preferred over FoundationPose 67% of the time. Ships 500 human-verified trajectories deployed on a 22-DoF Sharpa Wave hand across 10 real tasks.
- It removes the capture device from the bottom of the pyramid. Where DexUMI needs a sensorized glove and teleop needs a rig, Do As I Do works from any RGB video β the most abundant (L1) data β and carries it through the retargeting bridge (L4) with physics verification (L5). This is the device-free extreme of the human-video fork.
- Dynamics-aware, not kinematic-only. Retargeting inside a physics sim with force perturbation and contact-transition penalties directly targets the weakness of naive humanβrobot mapping (dropped force / broken grasps) β the exact L4 failure mode the data pyramid Β§4 flags.
- Honest about the yield. The pipeline's quality filter is severe (see Β§4), which is itself the useful signal: internet video is abundant but most clips don't survive faithful reconstruction β quantifying how leaky the L1βL4 path really is.
flowchart LR
V[Monocular RGB video<br/>internet Β· ego Β· generated] --> H[HaWoR<br/>hand tracking]
V --> O[SAM 3D<br/>object mesh + pose]
O --> T[Guided-diffusion object tracker<br/>fix shape @ anchor, track pose via flow matching]
H --> A[Align hand+object<br/>centroid opt + GeoCalib gravity]
T --> A
A --> R[Dynamics-aware retargeting<br/>MPPI-style optim in MuJoCo Warp<br/>+ warmup Β· force perturbation Β· transition reward]
R --> D[Robot-executable trajectory<br/>22-DoF Sharpa Wave + dual UR3e @ 50 Hz]
- Reconstruction: HaWoR hands + SAM-3D object geometry; the object tracker holds shape fixed at an anchor frame and recovers per-frame pose entirely at inference via flow matching with adaptive guidance; hand and object are aligned across scales by centroid optimization and gravity alignment.
- Retargeting: sampling-based optimization in MuJoCo Warp (200 Hz sim). Three robustifications for noisy references: warmup (hold the object while the hand settles β the single largest gain, 0.25β0.66 on reconstructed data), random force perturbation (encourages robust grasps), and a transition reward penalizing failed contact transitions.
- Hardware: 22-DoF Sharpa Wave hand on dual UR3e arms, commanded at 50 Hz.
Retargeting success (all components vs annealed sampling alone):
| Source | Baseline | Full method |
|---|---|---|
| Reconstructed in-the-wild video | 25% | 71% |
| OakInk2 (clean MoCap) | 72% | 81% |
- Reconstruction quality: on 150 in-the-wild videos, human raters prefer the object tracking over FoundationPose 67% of the time; SOTA F-5/F-10/Chamfer on DexYCB/HOI4D.
- Real-world: 10 tasks β whisking, pouring, dusting, squeezing, tamping, erasing, stirring, hammering, spreading, picking.
- Dataset produced: 500 high-quality, human-verified trajectories β 53% internet, 31% egocentric, 16% generated video.
- Yield reality-check: from 2,000 100DOH clips, only 83 (4%) survived the reconstruction pass before the final 500 validated trajectories were assembled.
Significance. Do As I Do is the strongest 2026 demonstration that device-free everyday video can become deployable dexterous-hand data, and its physics-in-the-loop retargeting is a concrete recipe for the pyramid's load-bearing L4 bridge. It also quantifies the leakiness of the human-video base (4% survival), which most L1-scaling papers gloss over.
Limitations (authors').
- Assumes rigid objects and semi-accurate monocular metric depth β fails otherwise.
- Monocular hand-object distance ambiguity is inherent.
- No environmental reasoning β obstacles, articulated objects out of scope.
- Sim is an upper bound β physics approximation caps real-world transfer.
- Low yield β heavy quality filtering means raw internet scale β usable-trajectory scale.
- Paper: arXiv 2606.19333
- Pyramid placement: L1 (web/ego video) β L4 (dynamics-aware retarget) β L5 (sim verify) β Dexterous-Hand Data Pyramid
- Counterpart bridge: AnyDexRT (calibration-free retargeting) Β· glove alternative: DexUMI Β· scaling law: EgoScale
- Dexterous Manipulation Β· Human Video β Robot Transfer
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)