arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16841cs.CVcs.MM

回答前看清楚:通过显著性驱动的感知重新对齐减轻LVLMs中的幻觉

Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment

Pengxu Chen, Yao Zhu, Guangming Zhu, Jun Sheng, Jincai Huang, Xiangyang Ji, Liang Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对LVLMs易产生幻觉问题,提出无需训练的SDPR框架,通过显著性驱动注意力重新分配、缓存对齐及先验约束对比解码,整体对齐视觉意识,在多基准测试中优于现有方法,无需额外训练且开销小。

中文摘要 AI 辅助

大型视觉语言模型(LVLMs)在多模态理解方面展现出卓越能力,但仍易产生与视觉证据不一致的幻觉。现有缓解方法多关注语言先验偏差或跨模态不平衡,而感知和记忆中的渐进视觉退化未被充分探索。本文提出显著性驱动的感知重新对齐(SDPR),这是一个无需训练的框架,可减轻推理过程中视觉意识的退化。具体包括:通过显著性驱动的注意力重新分配释放被非语义下沉令牌劫持的注意力,恢复关键视觉证据;识别KV缓存中的空间失真,提出显著性驱动的缓存对齐以在生成过程中保留与查询相关的视觉特征;引入先验约束对比解码以惩罚由主导语言先验引起的不忠实预测。大量实验表明,SDPR在幻觉和通用基准测试中均优于现有方法,无需额外训练且运行时开销最小。

英文摘要

Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding. However, they remain prone to hallucinations, generating responses that are inconsistent with the visual evidence. Existing mitigation methods largely address language-prior bias or cross-modal imbalance, while progressive visual degradation across perception and memory remains underexplored. In this work, we propose Saliency-Driven Perceptual Realignment (SDPR), a training-free framework that mitigates the degradation of visual awareness throughout inference. Specifically, we first introduce saliency-driven attention redistribution to release attention hijacked by non-semantic sink tokens, thereby recovering critical visual evidence. Second, we identify spatial distortion in the KV cache and propose saliency-driven cache alignment to preserve query-relevant visual features during generation. Finally, we introduce prior-constrained contrastive decoding to penalize unfaithful predictions induced by dominant language priors. Our proposed SDPR is robust against hallucinations due to its holistic alignment of visual awareness across the entire generative trajectory. Extensive experiments across diverse LVLM architectures show that SDPR outperforms state-of-the-art methods on both hallucination and general-purpose benchmarks, requiring no additional training and incurring minimal runtime overhead. The code is available \href{https://github.com/PengSyuChen/SDPR}{\color{blue}{here}}.

发表机构

  • Xidian University(西安电子科技大学)
  • Tsinghua University(清华大学)
  • Shanghai Road Transport Development Center(上海市道路运输发展中心)
  • Hunan Institute of Advanced Technology(湖南先进技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑