Skip to content

Review DoAsIDo

hwoo.han edited this page Aug 13, 2026 · 1 revision

In-Depth Review β€” Do As I Do: Dexterous Manipulation Data from Everyday Human Videos

Paper: "Do as I Do: Dexterous Manipulation Data from Everyday Human Videos" β€” arXiv 2606.19333 (Jun 17 2026) Authors: Bhawna Paliwal, Haritheja Etukuru, William Liang, Pieter Abbeel, Nur Muhammad Mahi Shafiullah, Jitendra Malik Β· UC Berkeley / NYU What it is: a device-free pipeline that turns ordinary monocular RGB human videos into robot-executable dexterous-hand trajectories β€” the cheapest L1β†’L4 path in the Dexterous-Hand Data Pyramid.

Companions: Dexterous-Hand Data Pyramid Β· AnyDexRT (the retargeting-bridge counterpart) Β· Dexterous Manipulation Β· Human Video β†’ Robot Transfer.


1. TL;DR

  1. Everyday video β†’ dexterous robot data, no gloves or teleop. Do As I Do reconstructs 4D hand-object dynamics from a single monocular RGB clip (internet, egocentric, or generated) and dynamics-aware retargets it into actions for a real dexterous hand.
  2. Reconstruction stack: HaWoR (hand tracking) + SAM 3D (object mesh/pose) + a guided-diffusion object tracker (fix shape at an anchor frame, track pose via flow matching at inference) + centroid/gravity alignment (GeoCalib).
  3. Dynamics-aware retargeting by sampling-based (MPPI-style) optimization in MuJoCo Warp, robustified for noisy references with warmup steps, random force perturbation, and a transition reward.
  4. Results: retargeting success 25% β†’ 71% on reconstructed in-the-wild video and 72% β†’ 81% on clean OakInk2 MoCap (vs annealed-sampling alone); object tracking preferred over FoundationPose 67% of the time. Ships 500 human-verified trajectories deployed on a 22-DoF Sharpa Wave hand across 10 real tasks.

2. Why it matters

  • It removes the capture device from the bottom of the pyramid. Where DexUMI needs a sensorized glove and teleop needs a rig, Do As I Do works from any RGB video β€” the most abundant (L1) data β€” and carries it through the retargeting bridge (L4) with physics verification (L5). This is the device-free extreme of the human-video fork.
  • Dynamics-aware, not kinematic-only. Retargeting inside a physics sim with force perturbation and contact-transition penalties directly targets the weakness of naive humanβ†’robot mapping (dropped force / broken grasps) β€” the exact L4 failure mode the data pyramid Β§4 flags.
  • Honest about the yield. The pipeline's quality filter is severe (see Β§4), which is itself the useful signal: internet video is abundant but most clips don't survive faithful reconstruction β€” quantifying how leaky the L1β†’L4 path really is.

3. Method

flowchart LR
  V[Monocular RGB video<br/>internet Β· ego Β· generated] --> H[HaWoR<br/>hand tracking]
  V --> O[SAM 3D<br/>object mesh + pose]
  O --> T[Guided-diffusion object tracker<br/>fix shape @ anchor, track pose via flow matching]
  H --> A[Align hand+object<br/>centroid opt + GeoCalib gravity]
  T --> A
  A --> R[Dynamics-aware retargeting<br/>MPPI-style optim in MuJoCo Warp<br/>+ warmup Β· force perturbation Β· transition reward]
  R --> D[Robot-executable trajectory<br/>22-DoF Sharpa Wave + dual UR3e @ 50 Hz]
Loading
  • Reconstruction: HaWoR hands + SAM-3D object geometry; the object tracker holds shape fixed at an anchor frame and recovers per-frame pose entirely at inference via flow matching with adaptive guidance; hand and object are aligned across scales by centroid optimization and gravity alignment.
  • Retargeting: sampling-based optimization in MuJoCo Warp (200 Hz sim). Three robustifications for noisy references: warmup (hold the object while the hand settles β€” the single largest gain, 0.25β†’0.66 on reconstructed data), random force perturbation (encourages robust grasps), and a transition reward penalizing failed contact transitions.
  • Hardware: 22-DoF Sharpa Wave hand on dual UR3e arms, commanded at 50 Hz.

4. Results (paper-reported)

Retargeting success (all components vs annealed sampling alone):

Source Baseline Full method
Reconstructed in-the-wild video 25% 71%
OakInk2 (clean MoCap) 72% 81%
  • Reconstruction quality: on 150 in-the-wild videos, human raters prefer the object tracking over FoundationPose 67% of the time; SOTA F-5/F-10/Chamfer on DexYCB/HOI4D.
  • Real-world: 10 tasks β€” whisking, pouring, dusting, squeezing, tamping, erasing, stirring, hammering, spreading, picking.
  • Dataset produced: 500 high-quality, human-verified trajectories β€” 53% internet, 31% egocentric, 16% generated video.
  • Yield reality-check: from 2,000 100DOH clips, only 83 (4%) survived the reconstruction pass before the final 500 validated trajectories were assembled.

5. Significance & limitations

Significance. Do As I Do is the strongest 2026 demonstration that device-free everyday video can become deployable dexterous-hand data, and its physics-in-the-loop retargeting is a concrete recipe for the pyramid's load-bearing L4 bridge. It also quantifies the leakiness of the human-video base (4% survival), which most L1-scaling papers gloss over.

Limitations (authors').

  1. Assumes rigid objects and semi-accurate monocular metric depth β€” fails otherwise.
  2. Monocular hand-object distance ambiguity is inherent.
  3. No environmental reasoning β€” obstacles, articulated objects out of scope.
  4. Sim is an upper bound β€” physics approximation caps real-world transfer.
  5. Low yield β€” heavy quality filtering means raw internet scale β‰  usable-trajectory scale.

6. Links

← Back to Reviews Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally