Skip to content

RSS 2026 LIBERO X

hwoo.han edited this page Jul 25, 2026 · 2 revisions

LIBERO-X β€” Robustness Litmus for Vision-Language-Action Models

Venue: RSS 2026 (Datasets & Benchmarks session) Β· Authors: Guodong Wang*, Chenkai Zhang*, Qingjie Liu, Jinjin Zhang, Jiancheng Cai, Junjie Liu, Xinmin Liu β€” Meituan Γ— Beihang University Β· arXiv: 2602.06556 Β· code Category: VLA benchmark Trend tag: RSS 2026 thread 6 β€” the evaluation crisis Scale: 5-level difficulty hierarchy Β· training set of 2,520 teleoperated demonstrations / 600 tasks / 100 scenes

Compiled from the verified RSS 2026 abstract and the paper's Fig. 1.

Key figure

Benchmark comparison (Figure 1 of arXiv 2602.06556, Β© the authors)

Figure 1 of the paper. (a) Original LIBERO: one scene β†’ one task β†’ homogeneous trajectories, with test data closely mirroring training β€” the memorization trap. (b) Current improvements (LIBERO-Plus/Pro): reuse the original trajectories and perturb single dimensions independently, with no systematic difficulty progression. (c) LIBERO-X: training becomes one scene β†’ multiple tasks β†’ diverse trajectories (the multi-colored trajectory bundle), and evaluation becomes a five-level pyramid β€” L1 local spatial perturbation β†’ L2 extended spatial perturbation β†’ L3 scene-topology reconstruction β†’ L4 visual-attribute variation β†’ L5 semantic-equivalent instruction reformulation β€” with the inset bar chart showing representative VLA success collapsing 39.4 β†’ 29.6 β†’ 17.0 β†’ 11.5 β†’ 8.2 as levels stack. The pyramid is the paper's argument: degradation is graded and diagnosable, not binary.

Problem

LIBERO-family benchmarks reward in-distribution pattern matching; existing protocols "provide limited or misleading assessments" because they don't capture real-world distribution shift β€” the same diagnosis as LIBERO-Plus and RobotManip Β§6.1.

Method

  1. Hierarchical evaluation protocol with progressive difficulty levels targeting three capabilities: spatial generalization, object recognition, and task-instruction understanding β€” enabling fine-grained analysis of where performance degrades as complexity compounds.
  2. High-diversity teleoperated training dataset in which each scene supports multiple fine-grained manipulation objectives, narrowing the train–eval distribution gap by construction rather than by perturbation alone.

Results (as reported)

  • Representative VLAs show significant performance drops under cumulative perturbations, exposing persistent weaknesses in scene comprehension and instruction grounding.

Significance

The third generation of the LIBERO robustness lineage: LIBERO β†’ LIBERO-Plus (seven perturbation axes) β†’ LIBERO-X (hierarchical capability-targeted protocol + train-side diversity). Its instruction-understanding axis converges with RoboTwin-IF's language-grounding diagnosis β€” the field is standardizing on capability-decomposed OOD evaluation. Useful companion to PolaRiS: LIBERO-X stresses a fixed sim benchmark correctly; PolaRiS makes new benchmarks from reality.

← RSS 2026 survey Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally