Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

3 Commits
 
 
 
 
 
 

Repository files navigation

ODI-Bench

arXiv Benchmark

Liu Yang1*, Huiyu Duan1*†, Ran Tao2, Juntao Cheng1, Sijing Wu1, Yunhao Li1, Jing Liu3, Xiongkuo Min1, Guangtao Zhai1

1Shanghai Jiao Tong University · 2Xinjiang University · 3Tianjin University

* Equal contribution. † Corresponding author.


Omnidirectional images (ODIs) provide full 360° × 180° view which are widely adopted in VR, AR and embodied intelligence applications. While multi-modal large language models (MLLMs) have demonstrated remarkable performance on conventional 2D image and video understanding benchmarks, their ability to comprehend the immersive environments captured by ODIs remains largely unexplored. To address this gap, we first present ODI-Bench, a novel comprehensive benchmark specifically designed for omnidirectional image understanding. We further introduce Omni-CoT, a training-free method which significantly enhances MLLMs’ comprehension ability in the omnidirectional environment through chain-of-thought reasoning across both textual information and visual cues.

Release


  • 2026-03-09 ❤️ The benchmark is currently released on Hugging Face!
  • 2026-01-26 🚀 Our paper ODI-BENCH is accepted by ICLR 2026! See u in Brazil!

If you find our work useful, please cite:

@article{yang2025odi,
 title={ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?},
 author={Yang, Liu and Duan, Huiyu and Tao, Ran and Cheng, Juntao and Wu, Sijing and Li, Yunhao and Liu, Jing and Min, Xiongkuo and Zhai, Guangtao},
 journal={arXiv preprint arXiv:2510.11549},
 year={2025}
}

About

[ICLR 2026] ODI-Bench: Can MLLMs understand immersive omnidirectional environments?

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors