-
Notifications
You must be signed in to change notification settings - Fork 0
Review VLA Evaluation
hwoo.han edited this page Aug 11, 2026
·
4 revisions
Topic survey (updated Aug 2026) Β· the evaluation crisis and its 2026 answers. Structure: π trend Β· βοΈ approaches Β· β emerging norms Β·
β οΈ limitations. Companion pages: PolaRiS Β· LIBERO-X Β· LIBERO-Plus Β· Qwen-RobotManip (the OOD manifesto) Β· WAM vs VLA Robustness Β· RoboMME.
Three independent 2026 demonstrations:
- From-scratch β pretrained, in-distribution. RobotManip Β§6.1: models with no large-scale robot pretraining match or beat Ο0.5-class models on LIBERO/RoboTwin (StarVLA 98.0, scratch 98.2 vs Ο0.5 97.6 on LIBERO) β the benchmarks reward memorization, not the pretraining the field spends its money on. The same paper's scaling ablation shows in-distribution scores are flat in pretraining data volume while OOD scores scale.
- Cliff-edge collapses under compound shift. LIBERO-X's L1βL5 pyramid: representative VLAs fall 39.4 β 29.6 β 17.0 β 11.5 β 8.2 as perturbations stack; StarVLA drops 85.7 β 10.6 on RoboTwin CleanβRand.
- Sim benchmarks don't rank real policies. Existing sim suites correlate weakly with real-world generalist performance β the gap PolaRiS was built to close (and quantifies via 600 real + 93k sim paired rollouts).
| Era | Practice | Failure mode |
|---|---|---|
| β€2024 | Single-suite success rates (LIBERO, Calvin) | Train/test distribution overlap |
| 2025 | Perturbation extensions (LIBERO-Plus/PRO, SafeLIBERO) | Single-axis, independent perturbations; no difficulty progression |
| H1 2026 | Evaluation methodology becomes a publishable contribution class: hierarchical capability-decomposed protocols, real-to-sim reconstruction, statistical rigor, safety-aware sim, benchmark-free world-model evaluation | Fragmentation; each lab ships its own protocol |
| Approach | Definition | Pros | Cons / limits |
|---|---|---|---|
| Static sim benchmarks | Fixed scenes/tasks, shared splits | Cheap, comparable, reproducible | Saturated; memorization-prone; insensitive to pretraining |
| Capability-decomposed perturbation suites | Progressive, labeled perturbation axes (LIBERO-X's spatialβtopologyβattributeβsemantic pyramid; LIBERO-Plus's 7 dims; RoboTwin-IF's language-grounding probes; RoboTwin-XE's embodiment swap) | Diagnostic β tells you which capability failed | Still sim; training-side diversity must co-evolve or the gap is unfair |
| Real-to-sim reconstruction | Scan real scenes β neural reconstruction β interactive sim evals (PolaRiS) | Anchored to reality; validated rank-correlation; scalable env creation; shareable hubs | Physics fidelity for contact; validation so far concentrated on Ο-family policies |
| Statistically rigorous real evaluation | Confidence-aware comparison beyond binary success (RSS #76); betting-based sequential evaluation (RSS #90); offline policy evaluation via discounted liveness (RSS #154) | Correct inferences from few rollouts | Doesn't reduce rollout cost by orders of magnitude |
| Safety-aware evaluation | Damage-aware simulation (OopsieVerse, RSS #98) | Measures what deployment actually risks | Young; no adoption yet |
| World-model evaluation | Roll policies out in a learned WM (Ctrl-World, WorldGym, Qwen-RobotWorld's stated direction; dWorldEval makes actions first-class tokens) | Unlimited, benchmark-drift-free | WM fidelity is the evaluator's own confound; language-actioned WMs can't consume continuous policy actions (dWorldEval is the first crack in this) |
| Tiered diagnostic benchmarks (ICML 2026) | LIBERO-Gen splits ID / compositional / domain-generalization tiers; VLA-Arena (170 tasks, L0βL2, safety/distractor/extrapolation axes) | Exposes spurious invariance standard metrics mask (Ο0.5: 64.0% Spatial-CG vs 21.2% Task-CG) | Still sim; tier taxonomies not yet standardized across suites |
| Failure diagnosis & adversarial probes (ICML 2026) | VLA-FixBench benchmarks 20 VLMs on fault diagnosis (idealized recovery loop: +13% LIBERO / +35% real); TRAP demonstrates the first targeted adversarial-patch attack on VLA chain-of-thought | Opens the recovery and security axes | Both nascent; no defense literature yet |
- Report OOD, not just in-distribution β with capability decomposition (spatial / object / instruction axes at minimum).
- Validate your sim against reality β PolaRiS-style paired-rollout correlation, not assumed transfer.
- Fine-tune-then-perturb protocols (train on Clean, evaluate on Rand) expose what generalist pretraining actually bought.
- Instruction-following probes (RoboTwin-IF-class) catch VLA-to-VA degradation β GR00T-N1.7's 16.6% there, despite respectable visual-perturbation scores, shows why this axis must be separate.
- Statistical reporting: trials-per-cell and uncertainty, especially for β€10-rollout real evaluations.
- Contact-rich and deformable-object tasks still lack faithful simulated evaluation anywhere β the shared blind spot of both perturbation suites and real-to-sim.
- Real-correlation evidence (PolaRiS) covers a narrow policy family so far; correlation for AR/tokenized or humanoid policies is unmeasured.
- Humanoid evaluation lags a full generation: intervention-assisted scoring, no OOD protocol, no perturbation suite (Review-Humanoid-VLA).
- No venue-level enforcement: OOD reporting remains voluntary, so the in-distribution leaderboard culture persists in parallel.
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)