|
|
|
|
|
DriveMLM |
 DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving |
arXiv 2023 |
- |
 |
RAG-Driver |
 RAG-Driver: Generalisable Driving Explanations with Retrieval-Augmented In-Context Learning in Multi-Modal Large Language Model |
RSS 2024 |
 |
 |
RDA-Driver |
 Making Large Language Models Better Planners with Reasoning-Decision Alignment |
ECCV 2024 |
- |
- |
DriveLM |
 DriveLM: Driving with Graph Visual Question Answering |
ECCV 2024 |
 |
 |
DriveGPT4 |
 DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model |
RA-L 2024 |
 |
- |
DriVLMe |
 DriVLMe: Enhancing LLM-based Autonomous Driving Agents with Embodied and Social Experience |
IROS 2024 |
 |
 |
LLaDA |
 Driving Everywhere with Large Language Model Policy Adaptation |
CVPR 2024 |
 |
 |
VLAAD |
 VLAAD: Vision and Language Assistant for Autonomous Driving |
WACVW 2024 |
- |
 |
OccLLaMA |
 OccLLaMA: A Unified Occupancy-Language-Action World Model for Understanding and Generation Tasks in Autonomous Driving |
arXiv 2024 |
 |
- |
Doe-1 |
 Doe-1: Closed-Loop Autonomous Driving with Large World Model |
arXiv 2024 |
 |
 |
LINGO-2 |
 LINGO-2: Driving with Natural Language |
- |
 |
- |
SafeAuto |
 SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation Models |
ICML 2025 |
- |
 |
OpenEMMA |
 OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving |
WACV 2025 |
- |
 |
ReasonPlan |
 ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving |
CoRL 2025 |
- |
 |
WKER |
 World Knowledge-Enhanced Reasoning Using Instruction-Guided Interactor in Autonomous Driving |
AAAI 2025 |
- |
- |
OmniDrive |
 OmniDrive: A Holistic LLM-Agent Framework for Autonomous Driving with 3D Perception, Reasoning and Planning |
CVPR 2025 |
- |
 |
S4-Driver |
 S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual Representation |
CVPR 2025 |
 |
- |
Occ-LLM |
 Occ-LLM: Enhancing Autonomous Driving with Occupancy-BasedLarge Language Models |
ICRA 2025 |
- |
- |
DriveBench |
 Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives |
ICCV 2025 |
 |
 |
FutureSightDrive |
 FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving |
NeurIPS 2025 |
 |
 |
ImpromptuVLA |
 Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models |
NeurIPS 2025 |
 |
 |
Sce2DriveX |
 Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning |
RA-L 2025 |
- |
- |
EMMA |
 EMMA: End-to-End Multimodal Model for Autonomous Driving |
TMLR 2025 |
 |
- |
DriveAgent-R1 |
 DriveAgent-R1: Advancing VLM-Based Autonomous Driving with Hybrid Thinking and Active Perception |
arXiv 2025 |
- |
- |
Drive-R1 |
 Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning |
arXiv 2025 |
- |
- |
FastDriveVLA |
 FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-Based Token Pruning |
arXiv 2025 |
- |
- |
WiseAD |
 WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model |
arXiv 2025 |
 |
 |
AutoDrive-RΒ² |
 AutoDrive-RΒ²: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving |
arXiv 2025 |
- |
- |
OmniReason |
 OmniReason: A Temporal-Guided Vision-Language-Action Framework for Autonomous Driving |
arXiv 2025 |
- |
- |
OpenREAD |
 OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic |
arXiv 2025 |
- |
 |
dVLM-AD |
 dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning |
arXiv 2025 |
- |
- |
PLA |
 A Unified Perception-Language-Action Framework for Adaptive Autonomous Driving |
arXiv 2025 |
- |
- |
AlphaDrive |
 AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning |
arXiv 2025 |
- |
 |
CoReVLA |
 CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine |
arXiv 2025 |
 |
 |
WAM-Diff |
 WAM-Diff: A Masked Diffusion VLA Framework with MoE and Online Reinforcement Learning for Autonomous Driving |
arXiv 2025 |
- |
 |
VLADriveBench |
 VLADriveBench: Evaluating CoT-Action Relationship in VLA for Autonomous Driving |
arXiv 2026 |
- |
- |
BLUE |
 BLUE: Toward Better Language Use in Efficient Vision-Language-Action Models for Autonomous Driving |
arXiv 2026 |
- |
- |
DriveMA |
 DriveMA: Rethinking Language Interfaces in Driving VLAs with One-Step Meta-Actions |
arXiv 2026 |
- |
- |
C-CoT |
 C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving |
arXiv 2026 |
- |
- |
MAGNIFIED |
 MAGNIFIED: RL Fine-tuning of Multimodal Large Language Models for Motion Planning |
arXiv 2026 |
- |
- |
DriveReward |
 DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving |
arXiv 2026 |
- |
- |
nuReasoning |
 nuReasoning: A Reasoning-Centric Dataset and Benchmark for Long-Tail Autonomous Driving |
arXiv 2026 |
- |
- |
Decision-Making |
 Decision-Making with Lightweight Confidence-Aware Language Model for Autonomous Driving |
arXiv 2026 |
- |
- |
Is |
 Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation in Autonomous Driving Models |
arXiv 2026 |
- |
- |
ReasonBreak |
 ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving |
arXiv 2026 |
- |
- |
Intend, |
 Intend, Reflect, Refine: An Adaptive Multimodal Reflection Framework for Autonomous Driving |
arXiv 2026 |
- |
- |
Judge, Then Drive |
 Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving |
arXiv 2026 |
- |
- |
EvoDrive |
 EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents |
arXiv 2026 |
- |
- |
Unifying |
 Unifying Language-Action Understanding and Generation for Autonomous Driving |
arXiv 2026 |
- |
- |
MindDriver |
 MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving |
arXiv 2026 |
- |
- |
HERMES |
 HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving |
arXiv 2026 |
- |
- |
Counterfactual VLA |
 Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning |
arXiv 2025 |
- |
- |
OmniDrive-R1 |
 OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving |
arXiv 2025 |
- |
- |
BeLLA |
 BeLLA: End-to-End Birds Eye View Large Language Assistant for Autonomous Driving |
arXiv 2025 |
- |
- |
Reasoning |
 Reasoning-aware Speculative Decoding for Efficient Vision-Language-Action Models in Autonomous Driving |
arXiv 2026 |
- |
- |
What |
 What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning |
arXiv 2026 |
- |
- |
Teaching |
 Teaching Vision-Language-Action Models What to See and Where to Look |
arXiv 2026 |
- |
- |
Deferred |
 Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs |
arXiv 2026 |
- |
- |
Depth |
 Depth-Wise Probing and Pruning of the Planning Token in a Driving Vision-Language-Action Model |
arXiv 2026 |
- |
- |
CritiqueDriveVLM |
 CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving |
arXiv 2026 |
- |
- |
FactorDrive |
 FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving |
arXiv 2026 |
- |
- |
XCoT |
 XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving |
arXiv 2026 |
- |
- |
MVPruner |
 MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving |
arXiv 2026 |
- |
- |
|
|
|
|
|