-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 MolmoSpaces
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: Datasets and Benchmarks Β· paper #91 Authors: Yejin Kim, Wilbert Pumacay, Omar Rayyan, Max Argus, Winson Han, Eli Vanderbilt, Jordi Salvador, Abhay Deshpande, Rose Hendrix, Snehal Jauhri, Shuo Liu, Nur Muhammad Mahi Shafiullah, Maya Guru, Ainaz Eftekhar, Karen Farley, Donovan Clay, Jiafei Duan, Arjun Guru, Piper Wolters, et al. (Allen Institute for AI and collaborators) arXiv: 2602.11337 Β· program page
Summary compiled from the arXiv paper (v2, titled "MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation"); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1: the ecosystem at a glance β 230K interior environments, 130k objects with asset files/text/scale-mass metadata, 42M grasps, and multiple robot embodiments, loadable into Isaac, ManiSkill, and MuJoCo with high-fidelity physics; top-right scatter shows the strong correlation between MolmoSpaces-Bench success and real-world success.
Evaluating generalist robot policies requires coverage of the long tail of scenes, objects, and instructions that physical evaluation cannot provide; existing benchmarks are near saturation, focus on short-horizon skills in single scenes, and many sim-to-real evaluation pipelines are proprietary or closed.
MolmoSpaces (Allen Institute for AI and collaborators; fully open-source) comprises four parts: MolmoSpaces-Scenes β over 230k indoor environments in five datasets (120 hand-crafted single-room MSCrafted scenes, 110k procedural MSProc houses, 110k MSProcObja scenes adding Objaverse objects, 110k MSMultiType diverse layouts, and MSTwin, a digital twin of the authors' real kitchen), physics-tuned for MuJoCo, IsaacSim, and ManiSkill; MolmoSpaces-Objects β 130k+ rigid and articulated models with semantic/physical metadata; MolmoSpaces-Grasp β 42M+ annotated 6-DoF grasps over 48k interactive objects; and MolmoSpaces-Bench β eight base tasks (navigate-to, pick, pick-and-place, pick-and-place-next-to, pick-and-place-color, open, close, open-door) with verified-solvable trials, success conditions, and dense rewards. Evaluations use a Franka FR3 in DROID configuration for rigid-body manipulation and a Rainbow RB-Y1 for navigation.
Zero-shot evaluation of open-source policies (Ο0, Ο0-FAST, Ο0.5 DROID joint-position variants, CAP contact-action policies, and navigation baselines like RING/DualVLN) shows newer policies outperform older ones (e.g., Close task up to 84% for the best policy; Pick tops out around 34% among Ο/CAP variants in Fig. 10). Benchmark pick results correlate with 752 real-world RoboArena pick episodes at Pearson R = 0.96 and Spearman Ο = 0.98 (RΒ² β 0.92 for object picking). Distributional analyses expose failure modes: with DROID-frequent prompt phrasing Ο0 comes within 1% of Ο0.5 versus a 14% gap otherwise; perturbing initial joint positions degrades Ο0.5 while lighting changes barely matter; occluding the wrist camera drops Ο0.5 to 2% success versus 20% for third-person-camera occlusion.
The largest open, simulator-agnostic evaluation substrate for generalist policies to date, and its diagnostic findings (prompt sensitivity, wrist-camera reliance) directly inform VLA evaluation methodology. Related threads: Review-VLA-Evaluation.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)