| 2026 |
arXiv |
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation |
 |
website |
| 2026 |
arXiv |
A Pragmatic VLA Foundation Model |
 |
website |
| 2026 |
arXiv |
TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers |
 |
--- |
| 2026 |
arXiv |
Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization |
 |
website |
| 2026 |
arXiv |
RoboReward: General-Purpose Vision-Language Reward Models for Robotics |
--- |
website 小参数专用模型胜过大VLM,“专用场景微调”比“通用机器人预训练” |
| 2025 |
arXiv |
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning |
 |
website |
| 2025 |
arXiv |
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training |
--- |
website VLA评价体系 |
| 2025 |
arXiv |
Galaxea open-world dataset and g0 dual-system vla model |
 |
website |
| 2025 |
arXiv |
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight |
 |
Model & dataset |
| 2025 |
arXiv |
Motus: A Unified Latent Action World Model |
 |
website |
| 2025 |
arXiv |
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations |
 |
website |
| 2025 |
arXiv |
SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead |
 |
website |
| 2025 |
arXiv |
Training-Time Action Conditioning for Efficient Real-Time Chunking |
--- |
--- |
| 2025 |
arXiv |
Reinforcing Action Policies by Prophesying |
--- |
website |
| 2025 |
arXiv |
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight |
 |
--- |
| 2025 |
arXiv |
WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation Control |
 |
website |
| 2025 |
arXiv |
RynnVLA-002: A Unified Vision-Language-Action and World Model |
 |
--- |
| 2025 |
arXiv |
πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models |
 |
website |
| 2025 |
arXiv |
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model |
 |
website |
| 2025 |
arXiv |
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models |
--- |
--- |
| 2025 |
arXiv |
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation |
--- |
--- |
| 2025 |
arXiv |
Libero-plus: In-depth robustness analysis of vision-language-action models |
 |
website |
| 2025 |
arXiv |
TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos |
--- |
website |
| 2025 |
arXiv |
VLA-Pruner: Temporal-Aware Dual-Level Visual Token Pruning for Efficient Vision-Language-Action Inference |
--- |
--- |
| 2025 |
arXiv |
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment |
 |
--- |
| 2025 |
arXiv |
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model |
 |
website |
| 2025 |
arXiv |
Evo-0: Vision-language-action model with implicit spatial understanding |
 |
website |
| 2025 |
arXiv |
EvoVLA: Self-Evolving Vision-Language-Action Model |
 |
website |
| 2025 |
arXiv |
Towards Deploying VLA without Fine-Tuning: Plug-and-Play Inference-Time VLA Policy Steering via Embodied Evolutionary Diffusion |
--- |
website |
| 2025 |
arXiv |
π0.6: a VLA That Learns From Experience |
--- |
website |
| 2025 |
ICCV |
Robomm: All-in-one multimodal large model for robotic manipulation |
 |
website |
| 2025 |
arXiv |
InternVLA-M1: Latent Spatial Grounding for Instruction-Following Robotic Manipulation |
 |
Website |
| 2025 |
CoRL |
π0.5: a Vision-Language-Action Model with Open-World Generalization |
 |
Blog PI0.5 |
| 2025 |
CoRR |
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots |
 |
website |
| 2025 |
arXiv |
AnywhereVLA: Language-Conditioned Exploration and Mobile Manipulation |
 |
website |
| 2025 |
arXiv |
Hi robot: Open-ended instruction following with hierarchical vision-language-action models |
--- |
Website |
| 2025 |
arXiv |
Fast: Efficient action tokenization for vision-language-action models |
 |
Website PI0-Fast |
| 2025 |
CoRL |
Dexvla: Vision-language model with plug-in diffusion expert for general robot control |
 |
website |
| 2025 |
ICML |
DiffusionVLA: Scaling Robot Foundation Models via Unified Diffusion and Autoregression |
--- |
website DiVLA |
| 2025 |
IROS |
Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems |
 |
website GO-1 |
| 2025 |
arXiv |
Fine-tuning vision-language-action models: Optimizing speed and success |
 |
Website OpenVLA-OFT |
| 2025 |
CoRL |
OpenVLA: An Open-Source Vision-Language-Action Model |
 |
Website |
| 2025 |
ICLR |
Latent action pretraining from videos |
 |
website LAPA |
| 2024 |
RSS |
Octo: An Open-Source Generalist Robot Policy |
 |
Website |
| 2024 |
CoRR |
π0: A Vision-Language-Action Flow Model for General Robot Control |
 |
Blog PI0 |
| 2024 |
arXiv |
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation |
--- |
website |
| 2024 |
ICLR |
Unleashing large-scale video generative pre-training for visual robot manipulation |
 |
website GR1 |
| 2024 |
CoRL |
Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation |
 |
website |
| 2023 |
CoRL |
Voxposer: Composable 3d value maps for robotic manipulation with language models |
 |
website |
| 2023 |
CoRL |
Rt-2: Vision-language-action models transfer web knowledge to robotic control |
--- |
Website |
| 2023 |
RSS |
Learning fine-grained bimanual manipulation with low-cost hardware |
 |
Website ALOHA/ACT |
| 2023 |
CoRL |
Do as i can, not as i say: Grounding language in robotic affordances |
--- |
website SayCan |
| 2022 |
arXiv |
Rt-1: Robotics transformer for real-world control at scale |
 |
website Blog |