arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18320cs.CLcs.LG

注意力分散作为大型语言模型幻觉的诊断信号

Attention Dispersion as a Diagnostic Signal for Hallucination in Large Language Models

Shardul P. More, Tanuja S. Pawar

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种基于注意力分散的无监督度量,用于检测大型语言模型在推理中的幻觉,在数学基准上相比输出置信度方法AUC提升最高达0.076。

中文摘要 AI 辅助

大型语言模型(LLMs)经常表现出幻觉,这成为复杂推理任务中可靠性的主要障碍。虽然传统检测方法依赖于基于输出的置信度指标,但这些逻辑值常常因现代对齐技术而校准不当。在本文中,我们研究内部注意力机制的时间波动性,作为不依赖于输出校准的幻觉替代诊断信号。通过引入一种无监督的注意力分散度量,我们表明认知不确定性在中间层留下可测量的痕迹,其中注意力熵的尖峰与推理中断相关联。我们在数学推理基准(GSM8K和MATH-500)上使用Qwen2.5模型系列(1.5B和3B参数)评估我们的方法,发现在所有测试条件下,与基于输出的基线相比,AUC统计显著提升,最高达+0.076。这些发现表明,注意力分散是传统幻觉检测方法的有力补充,需要在更广泛的模型系列和任务领域中进行进一步研究。

英文摘要

Large Language Models (LLMs) frequently exhibit hallucinations, presenting a major barrier to reliability in complex reasoning tasks. While traditional detection methods rely on output-based confidence metrics, these logits are often miscalibrated by modern alignment techniques. In this paper, we investigate the temporal volatility of internal attention mechanisms as an alternative diagnostic signal for hallucination that does not depend on output calibration. By introducing an unsupervised metric for attention dispersion, we show that epistemic uncertainty leaves a measurable trace within intermediate layers, where spikes in attention entropy are associated with reasoning breakdowns. We evaluate our approach on mathematical reasoning benchmarks (GSM8K and MATH-500) using the Qwen2.5 model family (1.5B and 3B parameters), finding statistically significant AUC improvements of up to +0.076 over output-based baselines across all tested conditions. These findings suggest that attention dispersion is a promising complement to traditional hallucination detection methods, requiring further investigation across broader model families and task domains.

发表机构

  • Rajarambapu Institute of Technology(拉贾兰巴普理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑