Repository navigation
ICLR 2026 X VLA
Venue: ICLR 2026 Category: VLA Architecture β Cross-Embodiment Trend tag: Trend 4
flowchart LR
D1[Dataset 1: Franka] --> SP1[Soft prompt 1<br/>few learned tokens]
D2[Dataset 2: UR5] --> SP2[Soft prompt 2]
D3[Dataset 3: dex hand] --> SP3[Soft prompt 3]
SP1 --> B[SHARED Transformer Backbone]
SP2 --> B
SP3 --> B
B --> A[Action]
Mixing robots with different dynamics, sensors, and action spaces during cross-embodiment training pulls a shared backbone in conflicting directions. The challenge is to exploit the heterogeneous cross-embodiment features rather than letting them interfere β converting potential conflict into positive cross-domain transfer while keeping the architecture simple and scalable.
A flow-matching VLA built exclusively on soft-prompted standard Transformer encoders (the 0.9B instantiation uses a 24-layer encoder, hidden size 1024, with a frozen Florence-Large VLM as the visionβlanguage encoder). Each data source / embodiment is assigned its own small set of learned soft-prompt embeddings, prepended to the input sequence (following the Lester et al. prompt-tuning recipe); the backbone is fully shared across embodiments, with only the soft prompt and the action-related input/output linear projections being embodiment-specific (together ~0.04% of total parameters). The prompt gives the backbone a clean "which robot am I?" signal so it can specialize internal computation without sacrificing shared representations. Actions are produced by an ODE-integrated flow-matching velocity field. Adaptation to a new robot is a two-phase recipe (Phase I generalist pretraining β Phase II domain adaptation): a fresh prompt is first warmed up with the backbone frozen, then optionally joint-finetuned; the PEFT variant tunes only 9M params (~1%) and matches Οβ on LIBERO/Simpler-WidowX despite ~300Γ fewer trainable params.
X-VLA-0.9B reaches state-of-the-art across a wide sweep β 6 simulation suites (LIBERO, SimplerEnv-WidowX, CALVIN, RoboTwin-2.0, VLABench, plus NAVSIM) and 3 real-world platforms (WidowX, the unseen-during-pretraining AIRBOT, and an Agilex bi-manual robot). Headline numbers include LIBERO β98% and CALVIN ABCβD β4.43 average sequence length (out of 5). The model also won 1st place (Champion) at the AgiBot World Challenge (IROS 2025). Joint multi-domain adaptation preserves and in some cases improves per-domain success vs. single-domain fine-tuning, evidencing positive transfer. Trained on 290K episodes from 7 data sources, the scaling trend (model size, data diversity, data volume) shows no sign of saturation.
Converts cross-embodiment scaling from an architecture problem into a data problem. Soft prompts are extremely lightweight (a few tokens per embodiment), so adding a new robot to an existing policy is cheap. Probably the cleanest scaling story in 2026 VLA research.
- ICLR 2026 listing
For how X-VLA compares to the full cross-embodiment landscape (soft-prompt vs. unified tokens vs. invariant latents vs. human-video bridging vs. morphology-aware vs. scale+prompt vs. world-model): Cross-Embodiment Training Review.
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)