Skip to content

ICRA 2026 VLA Practicality

Heungwoo edited this page Jun 1, 2026 · 1 revision

Rethinking VLA Practicality β€” Benchmark + Improved Baseline

Venue: ICRA 2026 Β· Authors: Wenxuan Song, Jiayi Chen, Xiaoquan Sun, Huashuo Lei, Yikai Qin, Wei Zhao, Pengxiang Ding, Han Zhao, Tongxin Wang, Pengxu Hou, Zhide Zhong, Haodong Yan, Donglin Wang, Jun Ma, Haoang Li Β· arXiv: 2602.22663 Category: Benchmark / robustness for VLA Trend tag: Practicality (latency Β· data Β· robustness)

Approach diagram

flowchart LR
  subgraph CEBench["CEBench benchmark"]
    SIM["14.4k sim trajectories<br/>36 tasks"] --> DR["domain randomization<br/>(clutter Β· lighting Β· texture Β· table height)"]
    REAL["1.6k real-world trajectories<br/>8 tasks"] --> DR
    EMB["embodiments:<br/>single-arm Β· bimanual Β· mobile bimanual"] --> DR
  end
  DR --> EVAL["evaluate VLAs<br/>(OpenVLA, RDT-1B, TinyVLA, ACT, DP, …)"]
  subgraph BASE["LLaVA-VLA baseline (0.5B)"]
    MV["multi-view images<br/>(1st + 3rd person)"] --> VLM["LLaVA-OneVision-0.5B"]
    PROP["proprioception tokenizer"] --> VLM
    VLM --> ACT["action chunks (size 5)<br/>hybrid nav + manip action space"]
  end
  EVAL --> BASE
Loading

Problem

VLAs are pushed toward ever-larger backbones, costly large-scale pre-training, and single-embodiment scopes β€” yet these choices are rarely interrogated against practical deployment. The paper organizes its study around three questions: Q1 how much performance actually depends on parameter scale (and which techniques let small models match large ones); Q2 whether pre-training is necessary for small models in a target scenario; and Q3 how to define a unified action space for cross-embodiment manipulation, covering fixed-base and mobile robots. The authors argue existing benchmarks do not jointly stress data efficiency, visual generalization (via domain randomization), and cross-embodiment deployability.

Method

CEBench is a benchmark spanning single-arm, bimanual, and mobile-bimanual embodiments in both simulation and the real world, built with domain randomization (clutter, random lighting, diverse textures, variable table heights) so that seen vs domain-randomized (DR) generalization can be measured directly. It collects 14.4k simulated trajectories across 36 tasks and 1.6k expert-curated real-world trajectories across 8 tasks. Evaluated policies include OpenVLA, RDT-1B, TinyVLA, RoboFlamingo, and generative/diffusion methods alongside ACT and Diffusion Policy baselines.

LLaVA-VLA is the improved baseline: a lightweight VLA built on a pre-trained LLaVA-OneVision-0.5B backbone. It takes multi-view images (first- and third-person, concatenated), adds a proprioception tokenizer, and emits action chunks (chunk size 5). A hybrid action space mixes direction and value tokens so a single policy can switch between navigation and manipulation β€” making it, per the authors, an end-to-end VLA for mobile manipulation. Training is two-stage: post-training then fine-tuning, with post-training on 8Γ— NVIDIA H100 and fine-tuning on a single NVIDIA 4090 β€” i.e., consumer-grade-GPU reach.

Results

  • CALVIN: first sub-task success 96.2%, comparable to the 97.4% of its 7B counterpart; last sub-task 50.6%. The 0.5B model thus tracks a 7B model on the entry task.
  • RoboTwin: 40.3% average success on seen tasks vs 28.6% under domain randomization β€” quantifying the seenβ†’DR generalization gap.
  • Real-world bimanual: 44.2% seen / 30.7% DR average success, on par with or above compared baselines.

(Numbers above are quoted from the paper; latency/throughput figures were not confirmable from the public text and are omitted.)

Significance

The paper reframes VLA progress as a practicality problem rather than a scale race: a 0.5B model with proprioception tokenization, action chunking, and a unified nav+manip action space can rival 7B policies on entry tasks while training on a single 4090. CEBench's explicit seen vs domain-randomized split makes it a natural companion to the robustness/benchmark thread β€” RobustVLA (perturbation robustness), LIBERO-Plus (factor-controlled generalization probing) β€” and the cross-embodiment design connects to the taxonomy in VLA Architectures review. For practitioners, the headline is that small + cross-embodiment + DR-tested is a viable axis distinct from raw parameter count.

Links

Related pages

← Back to ICRA-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally