arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理解语言模型为何产生幻觉:针对先验知识的推理测试

Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors

Yangfan Hu, Xuhan Tong, Haoyue Bai, Xi Ding, Shashank Muralidhar Bharadwaj, Siyang Cao, Robert Nowak, Jiawei Zhang

arXiv 2607.00447首次发表:更新:

发表机构

University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出推理错位假说,认为幻觉源于模型偏向统计显著的潜在关联而非遵循提示约束,并通过TrapQA测试集验证了两种偏差模式。

AI 中文摘要

大型语言模型经常产生违反提示级约束的幻觉答案。一个关键的诊断问题是这些失败是否反映了知识缺失,或者模型拥有相关信息但遵循了错误的推理路径。我们将这种现象研究为推理错位:提示支持的答案与统计显著的潜在关联所偏好的答案之间的不匹配。我们通过一个潜在关键任务模型形式化了这一观点,其中预训练频率不平衡可能导致捷径路径主导约束敏感路径并引发正推理损失。该框架预测了两种失败模式:实体消歧中的任务检索偏差和动作选择中的关键选择偏差。我们引入了TrapQA,一个受控的诊断测试平台,包含两个组件。ScientistQA通过补充事实探针测试相似科学家之间的消歧,而Real-Life Constrained QA测试在显著捷径下日常约束的遵循。我们的结果表明,幻觉可能源于有偏的潜在推理,而不仅仅是知识缺失。

英文摘要

Large language models often produce hallucinated answers that violate prompt-level constraints. A key diagnostic question is whether these failures reflect missing knowledge, or whether the model has the relevant information but follows the wrong inference path. We study this phenomenon as inference misalignment: a mismatch between the answer supported by the prompt and the answer favored by statistically salient latent associations. We formalize this view with a latent key-task model, in which pretraining-frequency imbalance can cause a shortcut path to dominate the constraint-sensitive path and induce positive inference loss. The framework predicts two failure modes: task-retrieval bias in entity disambiguation and key-selection bias in action choice. We introduce TrapQA, a controlled diagnostic testbed with two components. ScientistQA tests disambiguation among similar scientists with supplementary factual probes, while Real-Life Constrained QA tests everyday constraint following under salient shortcuts. Our results show that hallucination can arise from biased latent inference rather than absent knowledge alone.

CommentsFindings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑