-
Notifications
You must be signed in to change notification settings - Fork 0
ICML 2026 VLA Arena
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models β structured robustness diagnostics for VLA policies
Venue: ICML 2026 (Poster) Category: Benchmark Affiliations: Borong Zhang, Jiahao Li, Jiachen Shen, Yishuai Cai, Yuhao Zhang, Yuanpei Chen, Juntao Dai, Jiaming Ji, Yaodong Yang Traction (2026-06): 17 citations (arXiv)
Vision-Language-Action (VLA) models report high headline success rates, but those single-number scores conflate genuine generalization with memorization of training distributions. They also obscure whether failures stem from the task structure, the language command, or the visual observation. Existing benchmarks rarely separate these axes or stress-test safety behavior, so it is hard to tell why a model succeeds or fails, or whether one model's ranking advantage is robust or an artifact of a particular difficulty setting.
VLA-Arena is a comprehensive, open-source benchmark built around a structured task design organized along three orthogonal dimensions: task structure, language commands, and visual observations. Rather than a flat list of tasks, the benchmark spans 11 task suites grouped into four diagnostic categories:
- Safety β whether the policy respects safety constraints,
- Distractor β robustness to irrelevant objects/clutter,
- Extrapolation β generalization beyond the training distribution,
- Long Horizon β multi-step task completion.
In total there are 170 tasks at three difficulty levels (L0βL2). To probe robustness systematically, the framework applies graded language perturbations (W0βW4) and visual perturbations (V0βV4) as diagnostic tools, isolating which input modality drives a failure. The release includes scaled dataset variants (VLA-Arena-S/M/L), an evaluation toolchain, reference models, and a public leaderboard.
flowchart LR
A[Task suites: 11] --> B{4 categories}
B --> S[Safety]
B --> D[Distractor]
B --> E[Extrapolation]
B --> L[Long Horizon]
A --> C[170 tasks Β· L0βL2]
C --> P1[Language perturb. W0βW4]
C --> P2[Visual perturb. V0βV4]
P1 --> R[Robustness diagnostics]
P2 --> R
Evaluating state-of-the-art VLAs across this structured suite surfaces three consistent failure modes: "memorization over generalization, superficial visual perception, and a neglect of safety constraints." The authors further report that "model rank reversals across L0βL2 validate that each level provides non-redundant insights" β i.e., the relative ordering of models flips between difficulty levels, demonstrating that a single difficulty setting (and a single headline number) would give a misleading picture of comparative model quality.
VLA-Arena reframes VLA evaluation from a leaderboard chase into a diagnostic exercise. By factoring tasks along structure/language/vision and layering graded perturbations plus an explicit Safety category, it gives the community a shared, reproducible way to attribute failures and to expose brittleness that aggregate success rates hide. The fully open release β datasets, toolchain, models, and leaderboard β lowers the barrier for systematic, comparable robustness reporting across future VLA work.
- ICML 2026: https://icml.cc/virtual/2026/poster/61388
β Back to ICML-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)