发表机构
Graduate School of Informatics, University of Amsterdam; Amsterdam University Medical Center; ILLC, University of Amsterdam(阿姆斯特丹大学信息学研究生院; 阿姆斯特丹大学医学中心; 阿姆斯特丹大学逻辑、语言与计算研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型中 animacy 概念相关电路,构建受控数据集对四个开放权重模型进行电路发现,通过实验和消融表明存在处理 animacy 的因果机制,发现的电路定位性差且泛化有限,证实 animacy 概念特性。
AI 中文摘要
在书面语言中区分有生命和无生命的概念需要的不仅仅是浅层文本处理,因为它涉及识别复杂的选择限制和上下文线索,如动词 - 论元交互。然而,当前的大语言模型似乎有能力做到这一点。我们研究大语言模型这种对 animacy 敏感的行为是否可以追溯到一组局部的因果相关组件和连接。为此,我们构建了一个最小对的受控数据集,并对四个开放权重模型进行电路发现。通过深入实验和消融,我们表明这些模型中确实存在负责处理 animacy 的因果机制,从而发现了一个 animacy 电路。同时,与其他已知电路相比,这个电路的定位性较差,并且仅在模型和 animacy 任务中部分泛化,证实了 animacy 概念的分布式、上下文相关和某种程度上的渐变性质。
英文摘要
Distinguishing animate from inanimate concepts in written language requires more than shallow text processing, as it involves recognizing complex selectional constraints and contextual cues, such as verb-argument interactions. Yet, current large language models (LLMs) appear to be capable of doing it. We investigate whether this animacy-sensitive behavior of LLMs can be traced to a localized set of causally relevant components and connections. To do so, we construct a controlled dataset of minimal pairs and perform circuit discovery on four open-weight models. Through in-depth experiments and ablations, we show that a causal mechanism responsible for handling animacy in these models does exist, thus discovering an animacy circuit. At the same time, this circuit appears to be less localized compared to other known ones and generalizes only partially across models and animacy tasks, confirming the distributed, context-dependent, and somewhat graded nature of the animacy concept.