-
Notifications
You must be signed in to change notification settings - Fork 0
ICRA 2026 VLA Practicality
Venue: ICRA 2026 Β· Authors: Wenxuan Song, Jiayi Chen, Xiaoquan Sun, Huashuo Lei, Yikai Qin, Wei Zhao, Pengxiang Ding, Han Zhao, Tongxin Wang, Pengxu Hou, Zhide Zhong, Haodong Yan, Donglin Wang, Jun Ma, Haoang Li Β· arXiv: 2602.22663 Category: Benchmark / robustness for VLA Trend tag: Practicality (latency Β· data Β· robustness)
flowchart LR
subgraph CEBench["CEBench benchmark"]
SIM["14.4k sim trajectories<br/>36 tasks"] --> DR["domain randomization<br/>(clutter Β· lighting Β· texture Β· table height)"]
REAL["1.6k real-world trajectories<br/>8 tasks"] --> DR
EMB["embodiments:<br/>single-arm Β· bimanual Β· mobile bimanual"] --> DR
end
DR --> EVAL["evaluate VLAs<br/>(OpenVLA, RDT-1B, TinyVLA, ACT, DP, β¦)"]
subgraph BASE["LLaVA-VLA baseline (0.5B)"]
MV["multi-view images<br/>(1st + 3rd person)"] --> VLM["LLaVA-OneVision-0.5B"]
PROP["proprioception tokenizer"] --> VLM
VLM --> ACT["action chunks (size 5)<br/>hybrid nav + manip action space"]
end
EVAL --> BASE
VLAs are pushed toward ever-larger backbones, costly large-scale pre-training, and single-embodiment scopes β yet these choices are rarely interrogated against practical deployment. The paper organizes its study around three questions: Q1 how much performance actually depends on parameter scale (and which techniques let small models match large ones); Q2 whether pre-training is necessary for small models in a target scenario; and Q3 how to define a unified action space for cross-embodiment manipulation, covering fixed-base and mobile robots. The authors argue existing benchmarks do not jointly stress data efficiency, visual generalization (via domain randomization), and cross-embodiment deployability.
CEBench is a benchmark spanning single-arm, bimanual, and mobile-bimanual embodiments in both simulation and the real world, built with domain randomization (clutter, random lighting, diverse textures, variable table heights) so that seen vs domain-randomized (DR) generalization can be measured directly. It collects 14.4k simulated trajectories across 36 tasks and 1.6k expert-curated real-world trajectories across 8 tasks. Evaluated policies include OpenVLA, RDT-1B, TinyVLA, RoboFlamingo, and generative/diffusion methods alongside ACT and Diffusion Policy baselines.
LLaVA-VLA is the improved baseline: a lightweight VLA built on a pre-trained LLaVA-OneVision-0.5B backbone. It takes multi-view images (first- and third-person, concatenated), adds a proprioception tokenizer, and emits action chunks (chunk size 5). A hybrid action space mixes direction and value tokens so a single policy can switch between navigation and manipulation β making it, per the authors, an end-to-end VLA for mobile manipulation. Training is two-stage: post-training then fine-tuning, with post-training on 8Γ NVIDIA H100 and fine-tuning on a single NVIDIA 4090 β i.e., consumer-grade-GPU reach.
- CALVIN: first sub-task success 96.2%, comparable to the 97.4% of its 7B counterpart; last sub-task 50.6%. The 0.5B model thus tracks a 7B model on the entry task.
- RoboTwin: 40.3% average success on seen tasks vs 28.6% under domain randomization β quantifying the seenβDR generalization gap.
- Real-world bimanual: 44.2% seen / 30.7% DR average success, on par with or above compared baselines.
(Numbers above are quoted from the paper; latency/throughput figures were not confirmable from the public text and are omitted.)
The paper reframes VLA progress as a practicality problem rather than a scale race: a 0.5B model with proprioception tokenization, action chunking, and a unified nav+manip action space can rival 7B policies on entry tasks while training on a single 4090. CEBench's explicit seen vs domain-randomized split makes it a natural companion to the robustness/benchmark thread β RobustVLA (perturbation robustness), LIBERO-Plus (factor-controlled generalization probing) β and the cross-embodiment design connects to the taxonomy in VLA Architectures review. For practitioners, the headline is that small + cross-embodiment + DR-tested is a viable axis distinct from raw parameter count.
- arXiv: 2602.22663 Β· HTML Β· PDF
- RobustVLA (perturbation-robustness benchmark thread)
- LIBERO-Plus (factor-controlled generalization probing)
- VLA Architectures review (cross-embodiment / lightweight-backbone context)
- ICRA 2026 Survey
β Back to ICRA-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)