Skip to content

Review VLA Evaluation

hwoo.han edited this page Aug 11, 2026 · 4 revisions

Cross-Paper Review β€” Evaluating VLA Policies Credibly

Topic survey (updated Aug 2026) Β· the evaluation crisis and its 2026 answers. Structure: πŸ“ˆ trend Β· βš–οΈ approaches Β· βœ… emerging norms Β· ⚠️ limitations. Companion pages: PolaRiS Β· LIBERO-X Β· LIBERO-Plus Β· Qwen-RobotManip (the OOD manifesto) Β· WAM vs VLA Robustness Β· RoboMME.

1. The indictment β€” why in-distribution benchmarking broke

Three independent 2026 demonstrations:

  1. From-scratch β‰ˆ pretrained, in-distribution. RobotManip Β§6.1: models with no large-scale robot pretraining match or beat Ο€0.5-class models on LIBERO/RoboTwin (StarVLA 98.0, scratch 98.2 vs Ο€0.5 97.6 on LIBERO) β€” the benchmarks reward memorization, not the pretraining the field spends its money on. The same paper's scaling ablation shows in-distribution scores are flat in pretraining data volume while OOD scores scale.
  2. Cliff-edge collapses under compound shift. LIBERO-X's L1→L5 pyramid: representative VLAs fall 39.4 → 29.6 → 17.0 → 11.5 → 8.2 as perturbations stack; StarVLA drops 85.7 → 10.6 on RoboTwin Clean→Rand.
  3. Sim benchmarks don't rank real policies. Existing sim suites correlate weakly with real-world generalist performance β€” the gap PolaRiS was built to close (and quantifies via 600 real + 93k sim paired rollouts).

2. Trend arc

Era Practice Failure mode
≀2024 Single-suite success rates (LIBERO, Calvin) Train/test distribution overlap
2025 Perturbation extensions (LIBERO-Plus/PRO, SafeLIBERO) Single-axis, independent perturbations; no difficulty progression
H1 2026 Evaluation methodology becomes a publishable contribution class: hierarchical capability-decomposed protocols, real-to-sim reconstruction, statistical rigor, safety-aware sim, benchmark-free world-model evaluation Fragmentation; each lab ships its own protocol

3. Approach taxonomy (definitions, pros, cons)

Approach Definition Pros Cons / limits
Static sim benchmarks Fixed scenes/tasks, shared splits Cheap, comparable, reproducible Saturated; memorization-prone; insensitive to pretraining
Capability-decomposed perturbation suites Progressive, labeled perturbation axes (LIBERO-X's spatial→topology→attribute→semantic pyramid; LIBERO-Plus's 7 dims; RoboTwin-IF's language-grounding probes; RoboTwin-XE's embodiment swap) Diagnostic — tells you which capability failed Still sim; training-side diversity must co-evolve or the gap is unfair
Real-to-sim reconstruction Scan real scenes β†’ neural reconstruction β†’ interactive sim evals (PolaRiS) Anchored to reality; validated rank-correlation; scalable env creation; shareable hubs Physics fidelity for contact; validation so far concentrated on Ο€-family policies
Statistically rigorous real evaluation Confidence-aware comparison beyond binary success (RSS #76); betting-based sequential evaluation (RSS #90); offline policy evaluation via discounted liveness (RSS #154) Correct inferences from few rollouts Doesn't reduce rollout cost by orders of magnitude
Safety-aware evaluation Damage-aware simulation (OopsieVerse, RSS #98) Measures what deployment actually risks Young; no adoption yet
World-model evaluation Roll policies out in a learned WM (Ctrl-World, WorldGym, Qwen-RobotWorld's stated direction; dWorldEval makes actions first-class tokens) Unlimited, benchmark-drift-free WM fidelity is the evaluator's own confound; language-actioned WMs can't consume continuous policy actions (dWorldEval is the first crack in this)
Tiered diagnostic benchmarks (ICML 2026) LIBERO-Gen splits ID / compositional / domain-generalization tiers; VLA-Arena (170 tasks, L0–L2, safety/distractor/extrapolation axes) Exposes spurious invariance standard metrics mask (Ο€0.5: 64.0% Spatial-CG vs 21.2% Task-CG) Still sim; tier taxonomies not yet standardized across suites
Failure diagnosis & adversarial probes (ICML 2026) VLA-FixBench benchmarks 20 VLMs on fault diagnosis (idealized recovery loop: +13% LIBERO / +35% real); TRAP demonstrates the first targeted adversarial-patch attack on VLA chain-of-thought Opens the recovery and security axes Both nascent; no defense literature yet

4. Emerging norms worth adopting

  1. Report OOD, not just in-distribution β€” with capability decomposition (spatial / object / instruction axes at minimum).
  2. Validate your sim against reality β€” PolaRiS-style paired-rollout correlation, not assumed transfer.
  3. Fine-tune-then-perturb protocols (train on Clean, evaluate on Rand) expose what generalist pretraining actually bought.
  4. Instruction-following probes (RoboTwin-IF-class) catch VLA-to-VA degradation β€” GR00T-N1.7's 16.6% there, despite respectable visual-perturbation scores, shows why this axis must be separate.
  5. Statistical reporting: trials-per-cell and uncertainty, especially for ≀10-rollout real evaluations.

5. ⚠️ Limitations of the current toolkit

  • Contact-rich and deformable-object tasks still lack faithful simulated evaluation anywhere β€” the shared blind spot of both perturbation suites and real-to-sim.
  • Real-correlation evidence (PolaRiS) covers a narrow policy family so far; correlation for AR/tokenized or humanoid policies is unmeasured.
  • Humanoid evaluation lags a full generation: intervention-assisted scoring, no OOD protocol, no perturbation suite (Review-Humanoid-VLA).
  • No venue-level enforcement: OOD reporting remains voluntary, so the in-distribution leaderboard culture persists in parallel.

← Back to Home Β· Reviews

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally