超越注意力失衡:通过频谱手术缓解幻觉
Beyond Attention Imbalance: Mitigating Hallucinations via Spectral Surgery
浏览论文内容
中文总结 AI 辅助
针对LVLMs幻觉问题,本文提出无需训练和对比解码的FLASH框架,通过频谱涡旋分数检测视觉头并自适应调制频谱,纠正视觉信息流,实现性能与效率的更好平衡。
中文摘要 AI 辅助
尽管大型视觉-语言模型(LVLMs)取得了显著成功,幻觉仍然是其可靠部署的重大障碍。近期研究主要将这些问题的原因归咎于跨模态注意力失衡,因此大多数解决方案侧重于重新加权视觉令牌或抑制语言先验。然而,此类方法往往忽视了视觉信息流的频谱特性,并且常常依赖对比解码(CD),这会使推理时间加倍。我们没有遵循传统方法,而是识别出两种不同的幻觉模式——感知-语义分离和局部固着——并提出了FLASH(频率局部化注意力整形),一个无需训练且无需CD的框架。FLASH利用频谱涡旋分数来检测多头注意力层中的视觉头,并在解码过程中应用自适应频谱调制来纠正视觉信息流。实证结果表明,与最先进(SOTA)方法相比,FLASH在性能与效率之间实现了更优的平衡。
英文摘要
While Large Vision-Language Models (LVLMs) achieve remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these issues to cross-modal attention imbalances; most solutions therefore focus on reweighting visual tokens or suppressing language priors. However, such approaches often overlook the spectral characteristics of the visual information flow and frequently rely on Contrastive Decoding (CD), which doubles inference time. Instead of following conventional approaches, we identify two distinct hallucination patterns-Perceptual-Semantic Dissociation and Localized Fixation-and propose FLASH (Frequency-Localized Attention SHaping), a training-free and CD-free framework. FLASH utilizes a Spectral Vortex Score to detect vision heads within multi-head attention layers and applies adaptive spectral modulation to rectify the visual information flow during decoding. Empirical results demonstrate that FLASH achieves a superior balance between performance and efficiency compared to SOTA methods.
发表机构
- National University of Defense Technology(国防科技大学)
- Northwestern Polytechnical University(西北工业大学)
- National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean(空天海一体化国家工程实验室)
机构由 AI 辅助整理,请以论文原文为准。