arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15106cs.CL

当错误的键胜出:理解并检测大语言模型中的幻觉

When the Wrong Key Wins: Understanding and Detecting Hallucinations in LLMs

发表机构威斯康星大学麦迪逊分校
查看机构详情
  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

Xuhan Tong, Haoyue Bai, Dawei Zhou, Naichen Shi, Jiawei Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过潜在键视角揭示大语言模型幻觉源于预训练关联竞争,提出两阶段关键词扰动检测法,在多个基准上达0.910 AUROC,并统一解释四种幻觉机制。

中文摘要 AI 辅助

大型语言模型即使在正确回答所需的知识已经可用的情况下,也可能产生幻觉。我们通过推理的潜在键视角来研究这一失败,其中答案选择依赖于预训练期间习得的关联之间的竞争。我们表明,模型预测可能对单个查询关键词高度敏感,这些有影响力的关键词表现出实体特定的绑定,且其效应系统地受预训练频率的影响。多个绑定也可能在同一查询内竞争并表现出高阶交互。基于这一机制,我们提出了一种用于幻觉检测的两阶段关键词扰动方法。通过移除有影响力的关键词并测量模型如何重新组织其预测,该方法将误导性关键关联引起的错误与由诊断性证据支持的正确决策区分开来。在多个模型和基准测试中,扰动提供了强大且可迁移的检测信号,在探针已知的ScientistQA上达到0.910的AUROC。最后,我们将相同的概率框架扩展到四种幻觉机制:知识缺陷、错误知识、上下文干扰和不稳定推理。它们在基准测试中的操作分布为不同检测器家族为何在不同设置中成功提供了诊断背景。

英文摘要

Large language models can hallucinate even when the knowledge required for a correct answer is already available. We study this failure through a latent-key view of inference, where answer selection depends on competition among associations acquired during pretraining. We show that model predictions can be highly sensitive to individual query keywords, that these influential keywords exhibit entity-specific binding, and that their effects are systematically shaped by pretraining frequency. Multiple bindings can also compete and exhibit higher-order interactions within the same query. Based on this mechanism, we introduce a two-stage keyword-perturbation method for hallucination detection. By removing influential keywords and measuring how the model reorganizes its prediction, the method distinguishes errors caused by misleading key associations from correct decisions supported by diagnostic evidence. Across multiple models and benchmarks, perturbation provides a strong and transferable detection signal, reaching $0.910$ AUROC on probe-known ScientistQA. Finally, we extend the same probabilistic framework to four hallucination regimes: knowledge deficit, wrong knowledge, context distraction, and unstable inference. Their operational distributions across benchmarks provide diagnostic context for why different detector families succeed in different settings.

↑