✌️Hello! My name is Bill Wang (王文斌). I am a Ph.D. student in Artificial Intelligence at the School of Computer Science, Wuhan University. My current research focuses on image perception and reasoning for Multimodal Large Language Models.
🤔 I am interested in building AI systems that can perceive the visual world more precisely and reason about it more reliably. Despite recent progress in multimodal large language models (MLLMs), current AI systems still struggle with fine-grained visual details, high-resolution images, spatial reasoning in visual evidence. I believe that advancing visual intelligence requires more than recognizing objects or generating descriptions; AI systems should be able to identify critical visual evidence, understand complex visual structures, and reason over what they see.
👨🏻💻 To this end, my recent work explores spatial reasoning, high-resolution image perception, and AIGC detection. I have authored or co-authored more than ten papers in leading conferences and journals, including ICML, NeurIPS, CVPR, ACM MM, AAAI, Pattern Recognition. Previously, I interned at Tencent YouTu Lab and Hikvision.
🎯 My long-term goal is to develop AI systems with stronger, more reliable, and more efficient visual perception and reasoning capabilities.
📗 Publications: Checkout my publications on Google Scholar!


