通过跨模态注意力漂移与基于掩码的验证检测大型视觉语言模型中的对象幻觉
Detecting Object Hallucinations in Large Vision-Language Models via Cross-Modal Attention Drifts and Mask-Based Verification
浏览论文内容
中文总结 AI 辅助
针对大型视觉语言模型的对象幻觉问题,提出结合跨模态注意力漂移与掩码验证的CADMP框架,经多基准实验验证其检测性能具竞争力。
中文摘要 AI 辅助
尽管大型视觉语言模型(LVLMs)近期取得了进展,但对象幻觉仍是阻碍其可靠部署的主要障碍。现有检测方法常利用单一层的注意力来表征视觉 grounding,却未充分探索其跨层演化。本文提出CADMP,一种轻量级对象幻觉检测框架,结合相邻层跨模态注意力漂移与针对目标视觉掩码的预测敏感性。在解码过程中,CADMP量化连续跨模态注意力图间的分布变化,以捕捉视觉 grounding 的突变;随后选择漂移最大的突变,定位对应视觉相关区域,并测量掩码这些区域后预测概率的变化。这两种信号提供互补证据:注意力漂移表征内部视觉 grounding 的稳定性,概率变化则验证预测是否真正依赖于所识别的视觉证据。一个轻量级检测器整合两种信号以识别幻觉预测。在多个基准及代表性开源LVLMs上的实验表明,CADMP实现了始终具竞争力的检测性能;消融研究进一步证实,相邻层漂移建模与基于掩码的视觉 grounding 验证具有互补贡献。
英文摘要
Despite recent advances in large vision-language models (LVLMs), object hallucination remains a major barrier to their reliable deployment. Existing detection methods often characterize visual grounding using attention from individual layers, leaving its evolution across layers underexplored. We propose CADMP, a lightweight object hallucination detection framework that combines adjacent-layer cross-modal attention drift with prediction sensitivity to targeted visual masking. During decoding, CADMP quantifies distributional changes between consecutive cross-modal attention maps to capture abrupt transitions in visual grounding. It then selects the transition with the largest drift, locates the corresponding visually relevant regions, and measures the change in prediction probability after masking these regions. These two signals provide complementary evidence: attention drift characterizes the stability of internal visual grounding, while probability variation verifies whether a prediction truly depends on the identified visual evidence. A lightweight detector integrates both signals to identify hallucinated predictions. Experiments on multiple benchmarks and representative open-source LVLMs demonstrate that CADMP achieves consistently competitive detection performance. Ablation studies further confirm the complementary contributions of adjacent-layer drift modeling and mask-based grounding verification.
发表机构
- Zhongguancun Laboratory(中关村实验室)
机构由 AI 辅助整理,请以论文原文为准。