-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 ViserDex
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: RL Β· paper #150 Authors: Arjun Bhardwaj, Maximum Wilder-Smith, Mayank Mittal, Vaishakh Patil, Marco Hutter arXiv: 2604.11138 Β· program page
Summary compiled from the arXiv paper (v1, 13 Apr 2026); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Fig. 1: A scanned object becomes an augmentable 3D Gaussian scene that is composited into physics simulation and rendered via parallelized rasterization; training proceeds through privileged teacher training, teacherβstudent distillation, and pose-estimator training. The deployment strip shows a real multi-fingered hand reorienting the object under diverse (including adversarial) lighting and across multiple objects.
In-hand object reorientation requires precise object-pose estimation. RGB sensing offers rich semantic cues but existing solutions depend on multi-camera rigs or costly ray tracing, and generating photorealistic visual diversity for RGB sim-to-real via standard mesh rendering is computationally intractable, often demanding large compute clusters even for simple objects.
ViserDex is a monocular-RGB sim-to-real framework that integrates 3D Gaussian Splatting (3DGS) to bridge the visual gap. Its key insight is domain randomization directly in the Gaussian representation space: physically consistent pre-rendering augmentations (e.g. perturbations of spherical-harmonic coefficients) produce photorealistic, randomized data for pose estimation. The manipulation policy is trained with curriculum-based RL and teacherβstudent distillation, and both the perception and control models can be trained independently on consumer-grade hardware.
The pose estimator trained on 3DGS data reaches a mean accuracy of 65.4% and mean ADD of 10.2 mm, surpassing standard tiled rendering (53.3%) and a randomized-tiled baseline (55.6%). Under adversarial lighting it stays most robust at 56.3% mean accuracy versus tiled rendering 47.2% and naive GS 36.5%. The pre-rasterization augmentations add only β4% to frame-rendering time. On a physical multi-fingered hand with a single RGB camera, the system robustly reorients five diverse objects (over 25 consecutive reorientations) even under challenging lighting.
Demonstrates Gaussian splatting as a practical, low-compute path to RGB-only dexterous manipulation. Connects to RL and Review-Dexterous-Manipulation.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)