发表机构
KAIST; National Taiwan University of Science and Technology(韩国科学技术院; 台湾科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对10个LLMs,发现其处理信念与事实的缺陷程度取决于表达信念的动词,源于任务混淆,一条指令可扭转该失败,解码时抑制注意力仅能部分恢复准确率。
AI 中文摘要
人类在日常交流中会自然形成并表达信念,例如“我认为答案是3”或“我推测那是对的”。这类信念不可避免地与事实和知识交织,因此大型语言模型(LLMs)若要应用于面向用户的场景,同时处理信念与事实的能力就显得尤为重要。已有研究表明,即使是能力较强的LLMs,在承认基于错误信息的用户信念方面也存在系统性缺陷。我们将这一评估扩展至10个LLMs,涉及18种认知表达,发现该缺陷的规模和方向取决于表达信念所用的动词,事实信息与错误信息之间的准确率差距在“我模糊记得”时为+50%,在“我严重怀疑”时为-14%。我们进一步表明,该现象源于任务混淆:模型默认对潜在主张进行事实核查,从而覆盖了用户明确表达的信念;经过明确事实核查的思维链在错误信息上的准确率低于未进行此类核查的思维链;一条指令即可扭转不同动词类别的失败情况。从机制上看,模型会更多关注无法确认的错误信念,但在解码时抑制这种注意力仅能部分恢复准确率,且仅适用于部分模型,这需要未来研究干预方法。我们的发现厘清了先前的结果,并表明通常有益的事实核查行为会如何干扰LLMs的信念追踪。我们的代码可在该https URL获取。
英文摘要
Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevitably intertwine with fact and knowledge, making the ability to handle them in tandem desirable for large language models (LLMs), as they are increasingly deployed in user-facing settings. Prior work showed that even capable LLMs exhibit a systemic weakness in acknowledging user beliefs grounded in incorrect information. We extend this evaluation to 10 LLMs across 18 epistemic expressions and find that the size and direction of this weakness depend on the verb used to express the belief, with the accuracy gap between factual and false information ranging from +50% on "I vaguely remember" to -14% on "I seriously doubt". We further show that the phenomenon stems from what we call task confusion: models default to fact-checking the underlying claim, overriding the user's stated belief. We provide evidence where chains of thought that explicitly fact-check show lower accuracy on false information than those that do not, and a single instruction can reverse the failure across verb families. Mechanistically, models attend more to false beliefs they fail to confirm, but suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods. Our findings clarify prior results and show how fact-checking, a generally desirable behavior, can interfere with belief tracking in LLMs.
CommentsEMNLP 2026 (Main)