Skip to content
View DreamMr's full-sized avatar
💭
I may be slow to respond.
💭
I may be slow to respond.

Block or report DreamMr

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
DreamMr/README.md

About Me

✌️Hello! My name is Bill Wang (王文斌). I am a Ph.D. student in Artificial Intelligence at the School of Computer Science, Wuhan University. My current research focuses on image perception and reasoning for Multimodal Large Language Models.

🤔 I am interested in building AI systems that can perceive the visual world more precisely and reason about it more reliably. Despite recent progress in multimodal large language models (MLLMs), current AI systems still struggle with fine-grained visual details, high-resolution images, spatial reasoning in visual evidence. I believe that advancing visual intelligence requires more than recognizing objects or generating descriptions; AI systems should be able to identify critical visual evidence, understand complex visual structures, and reason over what they see.

👨🏻‍💻 To this end, my recent work explores spatial reasoning, high-resolution image perception, and AIGC detection. I have authored or co-authored more than ten papers in leading conferences and journals, including ICML, NeurIPS, CVPR, ACM MM, AAAI, Pattern Recognition. Previously, I interned at Tencent YouTu Lab and Hikvision.

🎯 My long-term goal is to develop AI systems with stronger, more reliable, and more efficient visual perception and reasoning capabilities.

📗 Publications: Checkout my publications on Google Scholar!

Popular repositories Loading

  1. RAP RAP Public

    Code for Retrieval-Augmented Perception (ICML 2025)

    Python 73 6

  2. HR-Bench HR-Bench Public

    PyTorch Implementation of "Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models"

    Python 49 2

  3. EST EST Public

    Expression Snippet Transformer for Robust Video-based Facial Expression Recognition

    Python 17 3

  4. WisdoM WisdoM Public

    Code for WisdoM: Improving Multimodal Sentiment Analysis by Fusing Contextual World Knowledge

    Jupyter Notebook 17

  5. TranX-Adapter TranX-Adapter Public

    Code for TranX-Adapter (ICML 2026)

    Python 15

  6. PCL PCL Public

    Pose-disentangled Contrastive Learning

    Python 14 5