About the problem formulation and evaluation results #7

MasterIzumi · 2024-04-19T19:57:38Z

Thanks for your great work!

I am wondering why GPT-Driver based on LLM outperforms the systems designed specifically for the motion planning task (e.g., NMP & UniAD, table I). If I understand correctly, it seems that you only provide limited information to LLM, while some important driving context is ignored (e.g., lane structure, traffic signals, etc.). Besides, when taking the visual grounding task as an example, although VLMs can detect and locate objects in the image, the accuracy can not be as good as the results obtained by object detectors. So in my personal view, LLMs/VLMs are not good at fine-grained spatial tasks at the current time point.

Can you provide some insight about the superior performance? Thanks!

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

About the problem formulation and evaluation results #7

About the problem formulation and evaluation results #7

MasterIzumi commented Apr 19, 2024

About the problem formulation and evaluation results #7

About the problem formulation and evaluation results #7

Comments

MasterIzumi commented Apr 19, 2024