-
Notifications
You must be signed in to change notification settings - Fork 0
CVPR 2026 FunREC
Venue: CVPR 2026 Category: 3D Scene Reconstruction for Manipulation Trend tag: Trend 5 Affiliations: ETH + MPI Informatics + Stanford + Microsoft + USI Lugano
flowchart LR
EGO["egocentric RGB-D video"] --> SEG["object segmentation"]
EGO --> KIN["kinematics<br/>extraction"]
SEG --> ARTIC["articulated-part decomposition"]
KIN --> ARTIC
ARTIC --> TWIN["functional 3D digital twin<br/>simulation-ready (URDF/USD)"]
TWIN --> POL["robot interaction<br/>(Spot mobile manipulator)"]
A digital twin for manipulation must include articulation (which parts move, how, with what kinematic constraints) β not just geometry. Building such twins by hand is expensive; learning them from video is hard because most video pipelines reconstruct only static geometry.
FunREC reconstructs simulation-ready functional 3D scenes with articulated parts and kinematics from egocentric RGB-D video. Operating on in-the-wild human interaction sequences (no controlled multi-state capture or CAD priors), the system automatically identifies articulated components, estimates their kinematic parameters along with per-timestep poses, and jointly reconstructs the static scene and each movable part (including interiors) in canonical space. Output is exported to URDF/USD for simulation.
Across two new benchmarks β RealFun4D (351 humanβscene interactions across 60 real apartments, captured with a head-mounted Azure Kinect DK) and OmniFun4D (127 photorealistic simulated interactions) β FunREC surpasses prior articulated-reconstruction work by a large margin: up to +50 mIoU in part segmentation, 5β10Γ lower articulation and pose errors, and substantially higher reconstruction accuracy.
For robotics, the reconstructed twins support URDF/USD export, hand-guided affordance mapping, and transfer of the human-demonstrated interaction to a Boston Dynamics Spot mobile manipulator using the inferred contact points and articulation parameters (demonstration-to-execution transfer, not a simulation-trained policy).
FunREC is the cleanest "digital twin from interaction" pipeline for manipulation to date. Closes a long-standing loop: ego video β articulated, simulation-ready digital twin (URDF/USD) β real-robot interaction. Note the robot result is demonstration-to-execution transfer on a Spot manipulator, not a closed-loop policy trained in the reconstructed sim. Closest predecessors: GenManip (LLM-driven scene graph sim) and RoboCasa365 (procedural assets at scale); FunREC's distinguishing feature is articulated-part inference from a single ego pass.
- arXiv: 2604.05621
- Project:
functionalscenes.github.io
β Back to CVPR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)