语法“祖母神经元”在大型语言模型中罕见
Grammatical "grandmother neurons" are rare in LLMs
浏览论文内容
中文总结 AI 辅助
本研究提出无探针框架及神经元可分性指数(NSI),在68个语言范式中发现LLM中强选择性“祖母神经元”罕见,且单神经元选择性与行为能力分离。
中文摘要 AI 辅助
理解大型语言模型(LLM)如何编码语言结构仍然是可解释性研究中的一个基本挑战。虽然诊断分类器(或“探针”)被广泛用于此任务,但它们面临重大的方法论批评:训练辅助分类器会引入容量混淆和校准问题,通常难以区分模型的内在表征与探针学习任务的能力。为解决这些局限,我们引入了一种无探针框架,用于在单个神经元层面定位语言选择性。利用语言最小对立的受控对比,我们提出了神经元可分性指数(NSI),该指标直接量化单个神经元在无参数更新的情况下区分语法与不合语法结构的可靠性。将NSI应用于68个语言范式和七个检查点,揭示了三个主要模式:1)对于形态和句法区分,原始可分性达到接近峰值水平的时间早于句法-语义接口和概念区分。2)经过排列归一化后,单单元选择性稀疏、微弱且窄调:只有一小部分单元对平均范式敏感,强选择性的“祖母神经元”罕见。3)全向量线性可分性、单神经元选择性和行为能力在很大程度上是分离的,定向消融进一步将激活选择性与因果依赖分开。
英文摘要
Understanding how Large Language Models (LLMs) encode linguistic structures remains a fundamental challenge in interpretability research. While diagnostic classifiers (or "probes") are widely used for this task, they face significant methodological criticism: training auxiliary classifiers introduces capacity confounds and calibration issues, often making it difficult to distinguish the model's intrinsic representations from the probe's ability to learn the task. To address these limitations, we introduce a probe-free framework for localizing linguistic selectivity at the individual neuron level. Leveraging the controlled contrasts of linguistic minimal pairs, we propose a Neuron Separability Index (NSI), a metric that directly quantifies how reliably single neurons differentiate grammatical from ungrammatical constructions without parameter updates. Applying NSI across 68 linguistic paradigms and seven checkpoints reveals three main patterns: 1) raw separability reaches near-peak levels earlier for morphological and syntactic distinctions than for syntax-semantics interface and conceptual distinctions. 2) after permutation normalization, single-unit selectivity is sparse, weak, and narrowly tuned: only a small fraction of units are sensitive to an average paradigm, and strongly selective "grandmother neurons" are rare. 3) whole-vector linear separability, single-neuron selectivity, and behavioral competence are largely dissociated, and targeted ablations further separate activation selectivity from causal reliance.
发表机构
- Zuckerman Mind Brain Behavior Institute, Columbia University(哥伦比亚大学祖克曼心智脑行为研究所)
机构由 AI 辅助整理,请以论文原文为准。