You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
DriveMA addresses the language–action gap in driving Vision-Language-Action models by introducing compact, verifiable meta-actions between visual observations and trajectory generation. It combines trajectory-grounded annotation, action-centric pretraining, and turn-level credit assignment reinforcement learning to make generated trajectories faithfully execute the model’s stated driving intent. DriveMA achieves state-of-the-art planning performance on WOD-E2E and competitive closed-loop results on NAVSIM.
Why is this technically valuable?
DriveMA makes intermediate language decisions verifiable by mapping predicted trajectories back into the meta-action space. Its turn-level RL assigns decision and trajectory rewards to their corresponding tokens, improving language–action consistency and enabling data-efficient, state-of-the-art driving performance.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Blog title
DriveMA: Driving Vision-Language-Action Models with Verifiable Meta-Actions
Canonical URL
https://tsinghua-mars-lab.github.io/DriveMA/
Are you the author?
Yes, I wrote this post
Author / Lab / Organization
Weicheng Zheng et al. — MARS Lab, IIIS, Tsinghua University
Source name
No response
Suggested BlogrXiv category
LLM & MLLM
Suggested tags
Autonomous Driving, VLA, Language-Action Alignment, RL
Short summary
DriveMA addresses the language–action gap in driving Vision-Language-Action models by introducing compact, verifiable meta-actions between visual observations and trajectory generation. It combines trajectory-grounded annotation, action-centric pretraining, and turn-level credit assignment reinforcement learning to make generated trajectories faithfully execute the model’s stated driving intent. DriveMA achieves state-of-the-art planning performance on WOD-E2E and competitive closed-loop results on NAVSIM.
Why is this technically valuable?
DriveMA makes intermediate language decisions verifiable by mapping predicted trajectories back into the meta-action space. Its turn-level RL assigns decision and trajectory rewards to their corresponding tokens, improving language–action consistency and enabling data-efficient, state-of-the-art driving performance.
Related paper / code, optional
https://github.com/Tsinghua-MARS-Lab/DriveMA
Submission confirmation
All reactions