发表机构
Universidade Estadual de Campinas (UNICAMP); University of Pennsylvania(坎皮纳斯州立大学; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过分析注意力图的拓扑结构(如Forman-Ricci曲率)检测大语言模型幻觉,提出捕捉半局部与全局信息流特征的单次通过方法,在基准上优于现有基线,并发现上下文共享受损是幻觉的关键成因。
AI 中文摘要
在本工作中,我们研究了注意力图内信息流模式的拓扑结构,以有效区分幻觉响应与非幻觉响应。我们分析了Forman-Ricci曲率,以识别指示注意力图中信息瓶颈的结构模式。随后,我们提出了一种方法,该方法捕捉与幻觉响应相关的注意力头的半局部和全局信息流特征。我们在多个大语言模型和既有基准上广泛评估了我们的方法。实证结果表明,我们提出的单次通过方法在两个幻觉检测基准上相较于现有的基于注意力和多响应的基线方法提供了持续改进,同时在多样化的大语言模型架构上取得了具有竞争力的性能。进一步的分析揭示,在因果生成过程中,token之间受损的上下文共享与幻觉发生密切相关。特别是,幻觉响应始终以过度依赖自注意力、从较早token的上下文检索分散或信息过度压缩为特征,尤其是在最后的Transformer层中。
英文摘要
In this work, we examine the topology of information flow patterns within attention graphs to effectively distinguish hallucinated from non-hallucinated responses. We analyze the Forman-Ricci curvature to identify structural patterns indicating information bottlenecks in attention graphs. We then introduce a method that captures both semi-local and global information-flow characteristics of attention heads associated with hallucinated responses. We evaluate our approach extensively across several LLMs and established benchmarks. Empirical results demonstrate that our proposed single-pass approach provides consistent improvements over existing attention-based and multi-response baselines across two hallucination-detection benchmarks, while achieving competitive performance across diverse LLM architectures. Further analysis reveals that impaired context sharing among tokens during causal generation is strongly associated with hallucination occurrences in LLMs. In particular, hallucinated responses are consistently characterized by an over-reliance on self-attention, diffused context retrieval from earlier tokens, or information over-squashing, especially in the final transformer layer.