arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

否定条件下,大语言模型的预测如何变化?

How Do LLMs Change Predictions Under Negation?

Jongwook Yoon, Jongwon Lim, Sungjib Lim, Woojin Cho, Yohan Jo

arXiv 2610.09571首次发表:更新:

发表机构

Graduate School of Data Science, Seoul National University; Department of Computer Science and Engineering, Seoul National University(首尔大学数据科学研究生院; 首尔大学计算机科学与工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过机制分析揭示大语言模型在否定条件下常重复原答案,其机制为抑制原答案而非利用其排除,并提出针对性训练目标以减少失败并保持通用能力。

AI 中文摘要

否定是人类语言的一个重要特征,然而大语言模型(LLMs)在处理否定时仍不可靠。我们在自己的否定基准上评估了近期开源的闭源的大语言模型,发现在37%-71%的情况下,它们在否定条件下会重复相同的答案(例如,对于“什么不是西班牙的首都?”回答“马德里”)。为理解和解决这一脆弱性,我们从机制上考察了模型在否定条件下的运作方式。我们的主要发现是,专门的注意力头和MLP神经元通过以下方式共同实现否定:(1)抑制对原始答案(如“马德里”)的检索,同时(2)在答案类别内提升一个偏好的候选答案(如“巴黎”)。这与人类否定处理的描述形成对比,在人类处理中,关于原始答案的信息有助于确定应排除的内容。此外,我们发现这种与人类处理的差异是否定失败的一个关键来源:模型的机制依赖于抑制原始答案,而非利用它来确定应排除的内容,因此当抑制过弱或对特定答案的偏向妨碍其选择替代答案时,模型可能重复原始答案。为解决模型否定机制中的这一弱点,我们提出了一种训练目标,要求对更自信的原始预测产生更大的答案偏好偏移,并表明它比标准微调基线在减少否定失败的同时,对通用能力的退化更少。综合来看,我们的结果展示了机制分析如何揭示语言能力失败的原因,并指导针对潜在局限性的训练。

英文摘要

Negation is an essential feature of human language, yet large language models (LLMs) remain unreliable in processing it. We evaluate recent open-source and closed-source LLMs on our negation benchmark and find that, in 37-71% of cases, they repeat the same answer under negation (e.g., "Madrid" for "What is not the capital of Spain?"). To understand and address this brittleness, we mechanistically examine how models operate under negation. Our main finding is that specialized attention heads and MLP neurons jointly implement negation by (1) suppressing retrieval of the original answer (e.g., "Madrid") while (2) promoting a favored candidate within the answer category (e.g., "Paris"). This contrasts with accounts of human negation processing, in which information about the original answer helps to determine what should be excluded. Furthermore, we find that this difference from human processing is a key source of negation failures: the model's mechanism relies on suppressing the original answer rather than using it to determine what to exclude, so the model can repeat the original answer when suppression is too weak or when a bias toward particular answers prevents it from selecting an alternative. To address this weakness in the model's negation mechanism, we propose a training objective that requires larger shifts in answer preference for more confident original predictions, and show that it reduces negation failures with less degradation of general capabilities than standard fine-tuning baselines. Together, our results demonstrate how mechanistic analysis can reveal why a linguistic capability fails and guide training that targets the underlying limitation.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑