Skip to content

PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models

Latest

Choose a tag to compare

@ceodspspectrum ceodspspectrum released this 25 Feb 23:50

Large Vision-Language Models (LVLMs) often hallucinate objects by effectively ignoring the image and instead relying on previously generated output (prelim) tokens. We quantify this phenomenon using mutual information analysis and propose PAS (Prelim Attention Score) — a lightweight, training-free signal computed from attention weights over prelim tokens. PAS requires no extra computational passes and operates during inference, achieving state-of-the-art object-hallucination detection across multiple models and datasets, enabling real-time filtering and intervention.