-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 LIBERO X
Venue: RSS 2026 (Datasets & Benchmarks session) Β· Authors: Guodong Wang*, Chenkai Zhang*, Qingjie Liu, Jinjin Zhang, Jiancheng Cai, Junjie Liu, Xinmin Liu β Meituan Γ Beihang University Β· arXiv: 2602.06556 Β· code Category: VLA benchmark Trend tag: RSS 2026 thread 6 β the evaluation crisis Scale: 5-level difficulty hierarchy Β· training set of 2,520 teleoperated demonstrations / 600 tasks / 100 scenes
Compiled from the verified RSS 2026 abstract and the paper's Fig. 1.

Figure 1 of the paper. (a) Original LIBERO: one scene β one task β homogeneous trajectories, with test data closely mirroring training β the memorization trap. (b) Current improvements (LIBERO-Plus/Pro): reuse the original trajectories and perturb single dimensions independently, with no systematic difficulty progression. (c) LIBERO-X: training becomes one scene β multiple tasks β diverse trajectories (the multi-colored trajectory bundle), and evaluation becomes a five-level pyramid β L1 local spatial perturbation β L2 extended spatial perturbation β L3 scene-topology reconstruction β L4 visual-attribute variation β L5 semantic-equivalent instruction reformulation β with the inset bar chart showing representative VLA success collapsing 39.4 β 29.6 β 17.0 β 11.5 β 8.2 as levels stack. The pyramid is the paper's argument: degradation is graded and diagnosable, not binary.
LIBERO-family benchmarks reward in-distribution pattern matching; existing protocols "provide limited or misleading assessments" because they don't capture real-world distribution shift β the same diagnosis as LIBERO-Plus and RobotManip Β§6.1.
- Hierarchical evaluation protocol with progressive difficulty levels targeting three capabilities: spatial generalization, object recognition, and task-instruction understanding β enabling fine-grained analysis of where performance degrades as complexity compounds.
- High-diversity teleoperated training dataset in which each scene supports multiple fine-grained manipulation objectives, narrowing the trainβeval distribution gap by construction rather than by perturbation alone.
- Representative VLAs show significant performance drops under cumulative perturbations, exposing persistent weaknesses in scene comprehension and instruction grounding.
The third generation of the LIBERO robustness lineage: LIBERO β LIBERO-Plus (seven perturbation axes) β LIBERO-X (hierarchical capability-targeted protocol + train-side diversity). Its instruction-understanding axis converges with RoboTwin-IF's language-grounding diagnosis β the field is standardizing on capability-decomposed OOD evaluation. Useful companion to PolaRiS: LIBERO-X stresses a fixed sim benchmark correctly; PolaRiS makes new benchmarks from reality.
β RSS 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)