arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态大语言模型在协同头信息分布漂移时产生幻觉

MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads

Meng'en Qin, Junye Chen, Jucheng Liu, Yinchen Liu, Youlu Xing, Song Wang, Ruize Han

arXiv 2609.09206首次发表:更新:

发表机构

Shenzhen University of Advanced Technology(深圳理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出HEAL方法,通过头级信息解缠与校准,识别并缓解多模态大语言模型在协同头信息分布漂移时产生的幻觉,实验证明其有效提升模型可信度。

AI 中文摘要

多模态大语言模型(MLLMs)常常面临幻觉问题,从而阻碍了其可靠的实际应用。现有的基于注意力的缓解方法主要依赖间接信号(如注意力权重),这些信号无法准确反映幻觉生成背后的实际信息转移。在本文中,我们提出了HEAL,即头级信息解缠与校准方法,用于识别和缓解幻觉。HEAL首先对多头输出进行因果噪声干预,以过滤掉因果冗余的头。随后,它通过反事实双重差分法解缠剩余头内的信息分布,将头分为四种类型。通过分析,我们观察到:当协同头中的信息分布偏离健康均衡时,幻觉就会发生,而这与模态特定头的数量或强度并无强相关。受此洞察启发,HEAL将动态信息校准因子注入协同头的值向量中,并主动调节视觉-语言依赖关系,将输出分布引向事实证据。大量实验表明,HEAL能有效减少多种MLLMs的幻觉,为增强模型可信度提供了一条简单且可解释的路径。

英文摘要

Multimodal Large Language Models (MLLMs) often struggle with hallucinations, thus hindering their reliable practical applications. Existing attention-based mitigation methods mainly rely on indirect signals (e.g., attention weights) that fail to accurately reflect the actual information shift underlying hallucination generation. In this paper, we propose HEAL, Head-lEvel information disentAnglement and caLibration for identifying and mitigating hallucinations. HEAL first employs causal noise intervention on multi-head outputs to filter out causally redundant heads. Subsequently, it disentangles information distribution within the remaining heads via the counterfactual Difference-in-Differences, categorizing heads into four types. Through analysis, we observe: hallucinations happen when information distribution drifts away from a healthy equilibrium in synergy heads, not strongly correlated with the quantity or strength of modality-specific heads. Motivated by this insight, HEAL injects dynamic information calibration factors into the value vectors of synergy heads and actively regulates visual-language dependencies, steering the output distribution towards factual evidence. Extensive experiments demonstrate that HEAL effectively reduces hallucinations across multiple MLLMs, offering a simple and interpretable pathway to enhance model trustworthiness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑