arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有否定线索都同等重要:词缀否定带来更好的否定理解

Not all Negation Cues are Equal: Affixal Negations Yield Better Negation Understanding

Tian Tan, Eduardo Blanco

arXiv 2609.13685首次发表:更新:

发表机构

University of Arizona(亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建大规模数据集NegCue,预训练多种语言模型,发现词缀否定对否定理解提升最大,而单词否定增益有限,且进一步预训练对LM和LLM均有效。

AI 中文摘要

否定对于语言模型(LMs)和大型语言模型(LLMs)来说仍然是一个长期存在的挑战。先前的工作主要集中于少数高频单词否定线索,如not和never,对更广泛的否定类型和现代LLMs的探索有限。为弥补这一空白,我们构建了NegCue,一个包含超过180万个样本的大规模数据集,涵盖单词、多词和词缀否定,拥有超过200种独特线索。我们进一步在NegCue上对仅编码器LM和LLM进行预训练,以研究不同否定类型如何影响否定理解。在五个下游基准上的实验表明,在相同训练规模下,否定类型对性能提升的贡献不均。特别是,词缀否定带来最大的改进,而常被研究的单词否定的增益则相对有限。此外,我们的结果表明,进一步预训练能提高LM和LLM的否定理解能力。

英文摘要

Negation remains a longstanding challenge for both language models (LMs) and large language models (LLMs). Prior work mainly focuses on a small set of high-frequency single-word negation cues, such as not and never, with limited exploration of broader negation types and modern LLMs. To address this gap, we construct NegCue, a large-scale dataset containing over 1.8M samples spanning single-word, multi-word, and affixal negation with more than 200 unique cues. We further pre-train both encoder-only LMs and LLMs on NegCue to investigate how different negation types affect negation understanding. Experiments on five downstream benchmarks show that negation types contribute unevenly to performance gains under the same training scale. In particular, affixal negation yields the largest improvements, while the gains from the commonly studied single-word negation remain modest. Moreover, our results demonstrate that further pre-training improves negation understanding for both LMs and LLMs.

CommentsAccepted to the EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑