Skip to content

ICLR 2026 ArtVIP

Heungwoo edited this page Jun 1, 2026 · 1 revision

ArtVIP β€” articulated digital-twin assets for robot learning

Venue: ICLR 2026 Β· Authors: Zhao Jin, Zhengping Che, Tao Li, Zhen Zhao, Kun Wu, Yuheng Zhang, Yinuo Zhao, Zehui Liu, Qiang Zhang, Xiaozhu Ju, Jing Tian, Yousong Xue, Jian Tang Β· Paper: arXiv 2506.04941 (Jun 2025) Β· Category: Dataset / asset library + benchmark Β· Trend tag: High-fidelity sim assets for sim-to-real

Approach diagram

flowchart LR
  Model[Professional 3D modelers<br/>unified standards] --> Mesh[Precise geometric meshes<br/>high-res PBR textures]
  Mesh --> Phys[Fine-tuned dynamic params<br/>enhanced joint-drive equation]
  Phys --> Mod[Embedded modular<br/>interaction behaviors]
  Mod --> Aff[Pixel-level<br/>affordance annotations]
  Aff --> Assets[206 articulated objects<br/>26 categories Β· USD format]
  Assets --> Isaac[Isaac Sim<br/>RTX renderer + GPU physics]
  Isaac --> Val[Validation:<br/>imitation learning + RL]
Loading

Problem

Robot learning increasingly relies on simulation, which demands high-quality digital assets to bridge the sim-to-real gap. Existing open-source articulated-object datasets suffer from insufficient visual realism and low physical fidelity, limiting their usefulness for training policies that transfer to the real world.

Method

ArtVIP is a fully open-source library of high-quality digital-twin articulated objects plus indoor-scene assets, built specifically for Isaac Sim (chosen for its RTX renderer and GPU-parallel physics over MuJoCo/Webots):

  • Visual realism via precise geometric meshes and high-resolution textures, rendered with Physically Based Rendering (PBR).
  • Physical fidelity via fine-tuned dynamic parameters and an enhanced joint-drive equation (position- and velocity-dependent stiffness/damping) over Isaac Sim's default; collision handled via convex hull / fine-tuned collision / convex decomposition as needed.
  • Embedded modular interaction behaviors packaged inside assets (avoiding per-object hand-written joint scripts).
  • Pixel-level affordance annotations.

Scale: 26 categories, 206 articulated-object assets (household items, large/small furniture, large/small appliances), shipped in USD format with production guidelines, plus digital-twin and fully interactive environments.

Results

Feature-map visualization and optical motion capture are used to quantitatively demonstrate ArtVIP's visual and physical fidelity. Applicability is validated across imitation learning and reinforcement learning experiments. (Specific success/error metrics omitted here pending the full tables.)

Significance

A standards-driven, fully open asset library that raises both the visual and physical bar for articulated objects, directly targeting sim-to-real transfer. The embedded modular interactions and pixel-level affordances make assets reusable across tasks without bespoke scripting β€” useful infrastructure for any Isaac-Sim-based manipulation pipeline.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally