-
Notifications
You must be signed in to change notification settings - Fork 0
CoRL 2026 MolmoBOT
Venue: CoRL 2026 (Austin, TX, Nov 9β12) Β· Allen Institute for AI (Ai2). Paper: arXiv 2603.16861. Representative of: the sim-only zero-shot-transfer thesis β enough simulation diversity β real-world manipulation with no real training data. Companions: World Models Β· Cross-Embodiment Β· CoRL 2026 survey.

Real robot data is expensive and task-specific fine-tuning is the norm. MolmoBOT asks whether simulation alone β if made large and diverse enough β can produce a manipulation policy that transfers zero-shot to physical robots, with no real training data and no real fine-tuning.
- MolmoBot-Engine β an open-source procedural pipeline that generates robots, tasks, and scenes in MolmoSpaces (200k+ pre-built houses) on top of the MuJoCo simulator. It randomizes layouts, 6-DoF object poses, lighting, textures/materials, friction/mass/joint damping, camera extrinsics, and injects action noise. Assets come from iTHOR + Objaverse (filtered for graspable, watertight colliders); an expert planner does phase-based grasp/place trajectories (RB-Y1 uses CuRobo).
- MolmoBot-Data β ~1.8M expert trajectories (~300M frames, ~5,817 robot-hours) across 94,300 unique environments, 11,400+ pickup assets and 9,400+ receptacles. Generated at ~1,024 episodes/GPU-hour (100Γ A100-80GB, ~4,500 GPU-hours total).
- Policies β MolmoBot (Molmo2 VLM + flow-matching action head); MolmoBot-Pi0 (Οβ architecture, for controlled comparison); MolmoBot-SPOC (lightweight, edge-deployable, RL-fine-tunable). Static (Franka FR3) and mobile (Rainbow Robotics RB-Y1) manipulation.
- Real tabletop pick-and-place (DROID setup, 4 environments, 120 trials, zero-shot): MolmoBot 79.2% vs Οβ.β -DROID 39.2% and MolmoBot-Pi0 46.7%.
- Single-camera real kitchen (30 tasks): MolmoBot-Img 86.6% vs Οβ.β zero-shot 63.3%.
- Sim pick-and-place (200 eps): MolmoBot-Img 67.0% vs Οβ.β fine-tuned 46.0%, Οβ.β zero-shot 20.0%.
- Mobile (RB-Y1, sim): door-open specialist 77.7%; multitask 70.2% door, 44.8% pick. Real door opening was hard β 2/9 successful openings, with underrepresented handle configs a key failure mode.
It is a strong existence proof that scale + diversity of synthetic data can beat real-data-fine-tuned baselines on real hardware for rigid/articulated pick-and-place, and it ships the data-generation engine openly.
Limitations (reviewer): confined to rigid-body + articulated tasks β no contact-rich (insertion), deformables, or fluids/granular media; mobile-manipulation real-world transfer is still fragile (door opening low, sensitive to under-sampled geometries); real evals cover only two platforms.
- arXiv 2603.16861
- Survey: CoRL 2026 Β· Related: World Models Β· Cross-Embodiment
β Back to CoRL 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)