Skip to content

RSS 2026 ViserDex

hwoo.han edited this page Aug 10, 2026 · 2 revisions

ViserDex: Visual Sim-to-Real for Robust Dexterous In-hand Reorientation

Venue: RSS 2026 (Sydney, Jul 13–17) Β· Session: RL Β· paper #150 Authors: Arjun Bhardwaj, Maximum Wilder-Smith, Mayank Mittal, Vaishakh Patil, Marco Hutter arXiv: 2604.11138 Β· program page

Summary compiled from the arXiv paper (v1, 13 Apr 2026); all numbers quoted from the paper. Trend context: RSS 2026 survey.

ViserDex 3DGS sim-to-real pipeline (Figure 1 of arXiv 2604.11138, Β© the authors)

Fig. 1: A scanned object becomes an augmentable 3D Gaussian scene that is composited into physics simulation and rendered via parallelized rasterization; training proceeds through privileged teacher training, teacher–student distillation, and pose-estimator training. The deployment strip shows a real multi-fingered hand reorienting the object under diverse (including adversarial) lighting and across multiple objects.

Problem

In-hand object reorientation requires precise object-pose estimation. RGB sensing offers rich semantic cues but existing solutions depend on multi-camera rigs or costly ray tracing, and generating photorealistic visual diversity for RGB sim-to-real via standard mesh rendering is computationally intractable, often demanding large compute clusters even for simple objects.

Method

ViserDex is a monocular-RGB sim-to-real framework that integrates 3D Gaussian Splatting (3DGS) to bridge the visual gap. Its key insight is domain randomization directly in the Gaussian representation space: physically consistent pre-rendering augmentations (e.g. perturbations of spherical-harmonic coefficients) produce photorealistic, randomized data for pose estimation. The manipulation policy is trained with curriculum-based RL and teacher–student distillation, and both the perception and control models can be trained independently on consumer-grade hardware.

Results

The pose estimator trained on 3DGS data reaches a mean accuracy of 65.4% and mean ADD of 10.2 mm, surpassing standard tiled rendering (53.3%) and a randomized-tiled baseline (55.6%). Under adversarial lighting it stays most robust at 56.3% mean accuracy versus tiled rendering 47.2% and naive GS 36.5%. The pre-rasterization augmentations add only β‰ˆ4% to frame-rendering time. On a physical multi-fingered hand with a single RGB camera, the system robustly reorients five diverse objects (over 25 consecutive reorientations) even under challenging lighting.

Significance

Demonstrates Gaussian splatting as a practical, low-compute path to RGB-only dexterous manipulation. Connects to RL and Review-Dexterous-Manipulation.

← Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally