Skip to content

Latest commit

 

History

111 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Awesome Visual-Language-Action (VLA)

This repository contains a curated list of resources addressing the VLA (Visual Language Action).

If you find some ignored papers, feel free to create pull requests, or open issues.

Contributions in any form to make this list more comprehensive are welcome.

If you find this repository useful, a simple star should be the best affirmation. 😊

Feel free to share this list with others!

Overview


  • For an in-depth summary of selected papers, see Link
Year Venue Paper Title Repository Note
2026 arXiv DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation Github stars website
2026 arXiv A Pragmatic VLA Foundation Model Github stars website
2026 arXiv TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers Github stars ---
2026 arXiv Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization Github stars website
2026 arXiv RoboReward: General-Purpose Vision-Language Reward Models for Robotics --- website
小参数专用模型胜过大VLM,“专用场景微调”比“通用机器人预训练”
2025 arXiv MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning Github stars website
2025 arXiv Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training --- website
VLA评价体系
2025 arXiv Galaxea open-world dataset and g0 dual-system vla model Github stars website
2025 arXiv Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight Github stars Model & dataset
2025 arXiv Motus: A Unified Latent Action World Model Github stars website
2025 arXiv XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Github stars website
2025 arXiv SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead Github stars website
2025 arXiv Training-Time Action Conditioning for Efficient Real-Time Chunking --- ---
2025 arXiv Reinforcing Action Policies by Prophesying --- website
2025 arXiv Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight Github stars ---
2025 arXiv WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation Control Github stars website
2025 arXiv RynnVLA-002: A Unified Vision-Language-Action and World Model Github stars ---
2025 arXiv πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models Github stars website
2025 arXiv X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model Github stars website
2025 arXiv RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models --- ---
2025 arXiv VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation --- ---
2025 arXiv Libero-plus: In-depth robustness analysis of vision-language-action models Github stars website
2025 arXiv TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos --- website
2025 arXiv VLA-Pruner: Temporal-Aware Dual-Level Visual Token Pruning for Efficient Vision-Language-Action Inference --- ---
2025 arXiv Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment Github stars ---
2025 arXiv Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model Github stars website
2025 arXiv Evo-0: Vision-language-action model with implicit spatial understanding Github stars website
2025 arXiv EvoVLA: Self-Evolving Vision-Language-Action Model Github stars website
2025 arXiv Towards Deploying VLA without Fine-Tuning: Plug-and-Play Inference-Time VLA Policy Steering via Embodied Evolutionary Diffusion --- website
2025 arXiv π0.6: a VLA That Learns From Experience --- website
2025 ICCV Robomm: All-in-one multimodal large model for robotic manipulation Github stars website
2025 arXiv InternVLA-M1: Latent Spatial Grounding for Instruction-Following Robotic Manipulation Github stars Website
2025 CoRL π0.5: a Vision-Language-Action Model with Open-World Generalization Github stars Blog
PI0.5
2025 CoRR GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Github stars website
2025 arXiv AnywhereVLA: Language-Conditioned Exploration and Mobile Manipulation Github stars website
2025 arXiv Hi robot: Open-ended instruction following with hierarchical vision-language-action models --- Website
2025 arXiv Fast: Efficient action tokenization for vision-language-action models Github stars Website
PI0-Fast
2025 CoRL Dexvla: Vision-language model with plug-in diffusion expert for general robot control Github stars website
2025 ICML DiffusionVLA: Scaling Robot Foundation Models via Unified Diffusion and Autoregression --- website
DiVLA
2025 IROS Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems Github stars website
GO-1
2025 arXiv Fine-tuning vision-language-action models: Optimizing speed and success Github stars Website
OpenVLA-OFT
2025 CoRL OpenVLA: An Open-Source Vision-Language-Action Model Github stars Website
2025 ICLR Latent action pretraining from videos Github stars website
LAPA
2024 RSS Octo: An Open-Source Generalist Robot Policy Github stars Website
2024 CoRR π0: A Vision-Language-Action Flow Model for General Robot Control Github stars Blog
PI0
2024 arXiv GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation --- website
2024 ICLR Unleashing large-scale video generative pre-training for visual robot manipulation Github stars website
GR1
2024 CoRL Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation Github stars website
2023 CoRL Voxposer: Composable 3d value maps for robotic manipulation with language models Github stars website
2023 CoRL Rt-2: Vision-language-action models transfer web knowledge to robotic control --- Website
2023 RSS Learning fine-grained bimanual manipulation with low-cost hardware Github stars Website
ALOHA/ACT
2023 CoRL Do as i can, not as i say: Grounding language in robotic affordances --- website
SayCan
2022 arXiv Rt-1: Robotics transformer for real-world control at scale Github stars website
Blog

Efficient-VLA

Year Venue Paper Title Repository Note
2025 arXiv Running VLAs at Real-time Speed Github stars Blog
2025 arXiv NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies --- ---
2025 arXiv Action-aware dynamic pruning for efficient vision-language-action manipulation --- ---
2025 arXiv Edgevla: Efficient vision-language-action models --- ---
2025 arXiv Hume: Introducing System-2 Thinking in Visual-Language-Action Model Github stars website
2025 CoRL Flower: Democratizing generalist robot policies with efficient vision-language-action flow policies Github stars
Github stars
website
2025 arXiv Accelerating vision-language-action model integrated with action chunking via parallel decoding --- ---
2025 arXiv Fine-tuning vision-language-action models: Optimizing speed and success Github stars website
2025 arXiv Nina: Normalizing flows in action. training vla models with normalizing flows Github stars ---
2025 ICCV Saliency-aware quantized imitation learning for efficient robotic control --- ---
2025 arXiv Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered --- website
2025 arXiv VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation Github stars website
2025 arXiv CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding Github stars website
2025 arXiv Omnisat: Compact action token, faster auto regression --- ---
2025 arXiv RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models --- demo
2025 arXiv KV-Efficient VLA: A Method of Speed up Vision Language Model with RNN-Gated Chunked KV Cache Github stars ---
2025 arXiv Fast ECoT: Efficient Embodied Chain-of-Thought via Thoughts Reuse --- ---
2025 arXiv Ttf-vla: Temporal token fusion via pixel-attention integration for vision-language-action models --- ---
2025 arXiv Unified Vision-Language-Action Model Github stars website
2025 ICML Otter: A vision-language-action model with text-aware visual feature extraction Github stars Website
2025 arXiv Sqap-vla: A synergistic quantization-aware pruning framework for high-performance vision-language-action models Github stars ---
2025 arXiv Fastdrivevla: Efficient end-to-end driving via plug-and-play reconstruction-based token pruning --- ---
2025 RAL TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation Github stars Website
2025 arXiv Smolvla: A vision-language-action model for affordable and efficient robotics --- website
2025 arXiv NORA: A Small Open-Sourced Generalist Vision-Language-Action Model for Embodied Tasks Github stars Website
2025 arXiv MoLe-VLA: Dynamic Layer-Skipping Vision-Language-Action Model via Mixture-of-Layers for Efficient Robot Manipulation Github stars Website
2025 arXiv EfficientVLA: Training‑Free Acceleration and Compression for Vision‑Language‑Action Models --- ---
2025 arXiv Think Twice, Act Once: Token‑Aware Compression and Action Reuse for Efficient Inference in Vision‑Language‑Action Models --- ---
2025 arXiv ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning --- Website
2025 arXiv OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation Github stars Website
2025 arXiv SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model Github stars Website
2025 arXiv Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning Github stars Website
2025 arXiv SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration --- ---
2025 arXiv VLA‑Cache: Towards Efficient Vision‑Language‑Action Model via Adaptive Token Caching in Robotic Manipulation Github stars ---
2025 arXiv SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning --- ---
2025 RSS FAST: Efficient Action Tokenization for Vision-Language-Action Models --- Website
2025 arXiv Real-Time Execution of Action Chunking Flow Policies --- Website
2025 arXiv VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting --- ---
2025 arXiv BitVLA: 1-Bit Vision-Language-Action Models for Robotics Manipulation Github stars ---
2025 arXiv SpecVLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance --- ---
2024 NeurIPS RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation Github stars Website
2024 NeurIPS DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution Github stars ---
2024 IROS From LLMs to Actions: Latent Codes as Bridges in Hierarchical Robot Control --- ---
2024 CoRL HIRT: Enhancing Robotic Control with Hierarchical Robot Transformers --- ---
2024 arXiv Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation --- Website
2024 CoRL Robotic Control via Embodied Chain-of-Thought Reasoning Github stars Website

Survey Paper

Year Venue Paper Title Repository Note
2025 arXiv An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges Github stars website
2025 arXiv A survey on efficient vision-language-action models Github stars website
2025 arXiv Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey Github stars Blog
2025 arXiv A Survey on Vision-Language-Action Models: An Action Tokenization Perspective Github stars ---
2025 arXiv Vision-language-action models: Concepts, progress, applications and challenges --- Blog
2025 arXiv Vision language action models in robotic manipulation: A systematic review --- ---
2025 arXiv Large vlm-based vision-language-action models for robotic manipulation: A survey Github stars ---
2025 Information Fusion Exploring embodied multimodal large models: Development, datasets, and future directions --- ---
2025 arXiv Pure Vision Language Action (VLA) Models: A Comprehensive Survey --- ---

Other Resources

About

Paper Survey for Visual Language Action

Resources

Stars

94 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors