Add in-depth cross-paper review of cross-embodiment VLA training
Covers 27 papers across ICLR 2026, CoRL 2025, and the π series, plus
key 2024-2026 precursors (OXE/RT-X, Octo, CrossFormer, HPT, RDT-1B,
UniAct, GR00T N1, Gemini Robotics 1.5) and the AnyBody benchmark.
Page structure:
- Why cross-embodiment matters; π0.7 UR5e zero-shot laundry (85.6%
progress, 80% SR) matching expert teleop, vs AnyBody showing novel-
morphology extrapolation still fails
- Seven-category taxonomy:
A. Unified tokenized action space (RT-X, Octo, RDT-1B, FAST)
B. Soft-prompt / body-conditioning (X-VLA, HPT)
C. Embodiment-invariant latents (UniVLA, UniAct, X-Sim, UniSkill,
TraceVLA, XR-1 UVMC)
D. Human video / wearable bridging (DexUMI, EgoDex, Visual-Imit-
Humanoid, ImMimic)
E. Morphology-aware architecture (CrossFormer, HPT stems, GR00T,
WholeBodyVLA)
F. Scale + prompt expansion (pi0.5 -> pi0.6 -> pi0.7, Gemini
Robotics 1.5)
G. World-model-mediated (DreamGen, Cosmos Policy, Ctrl-World)
- Per-paper deep-dives with arXiv links, mechanisms, pros/cons
- Cross-axis comparison (cheapest new-robot extension, most diverse
morphology span, most data-efficient, strongest zero-shot result)
- Trends: 2023 no-sharing -> 2024 Group A dominance -> 2025 Groups C
and D bloom -> late 2025/2026 Group F crowns production -> ICLR 2026
synthesizes B+C+G
- Open questions: does scale alone solve it; sim-to-real vs real-to-
sim; universal action space; strategy transfer; benchmark maturity;
open vs closed ecosystem
- Practical decision guide for shipping cross-embodiment VLAs
Cross-links added in X-VLA and X-Sim summary pages; sidebar and Home
updated to surface the new review.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>