发表机构
School of Artificial Intelligence, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Science; Hong Kong University of Science and Technology(中国科学院大学人工智能学院; 中国科学院自动化研究所; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LVLM的物体幻觉问题,通过Logit Lens分析视觉注意力,揭示两种幻觉机制,提出无需训练的Detect-Mitigate框架,在多个基准上实现最优结果。
AI 中文摘要
大型视觉语言模型(LVLMs)常受物体幻觉问题困扰,生成图像中不存在的物体。过往研究多将其归因于视觉注意力不足,但我们发现,在模型的中后层,真实物体和幻觉物体获得的视觉注意力强度相当,表明关键问题可能并非模型的注意力程度,而是其关注的内容及原因。为此,我们使用Logit Lens解码高注意力区域的视觉特征,观察到对应真实物体的区域可被正确解码为目标物体 token,而幻觉物体对应的区域则无法。基于此,我们确定两种幻觉机制:(i)视觉不确定性,由语义相似或易混淆的区域触发,屏蔽这些区域可消除幻觉;(ii)上下文先验,由强共现先验触发,即使初始关注区域被屏蔽,幻觉仍会持续且注意力会漂移到其他区域。基于这些发现,我们提出一种简单且无需训练的Detect-Mitigate框架,包含用于检测幻觉的Logit-Lens一致性检查,以及针对性补救措施:针对视觉不确定性幻觉的高注意力区域屏蔽(HARM),和针对上下文先验幻觉的视觉证据增强解码(VEED)。我们的方法在多个幻觉基准上取得了最优结果,代码将公开。
英文摘要
Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image. Prior work largely attributes this to insufficient visual attention. However, we find that both real and hallucinated objects receive equally strong visual attention in the model's mid-to-late layers, suggesting that the key issue may not be how much the model attends, but what it attends to and why. To this end, we decode the visual features of high-attention regions using Logit Lens, and observe that regions corresponding to real objects can be correctly decoded to the target object tokens, whereas those for hallucinated objects cannot. Building on this, we identify two hallucination mechanisms: (i) visual uncertainty, triggered by semantically similar or confusable regions; masking these regions eliminates the hallucination. (ii) contextual prior, triggered by strong co-occurrence priors; even when the initially attended region is masked, the hallucination persists and attention drifts to other regions. Based on these findings, we propose a simple yet effective training-free Detect-Mitigate framework comprising a Logit-Lens Consistency Check to detect hallucination and targeted remedies: High-Attention Regions Masking (HARM) for visual uncertainty hallucination, and Visual Evidence Enhanced Decoding (VEED) for contextual prior hallucination. Our approach achieves state-of-the-art results on multiple hallucination benchmarks. Code will be available.
CommentsCVPR2026 Highlight