发表机构
College of Biomedical Engineering, Fudan University(复旦大学生物医学工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对17个不同规模的LLMs审计4种神经科学启发范式,发现AI神经科学研究受测量混淆限制,发布了相关方案、刺激与代码。
AI 中文摘要
大型语言模型(LLMs)越来越多地被报道表现出类人的神经与认知特征,包括概念细胞、心理数字线和认知地图。这些结论通常依赖于对单一模型的线性探测和激活操控,但这两种方法对测量选择高度敏感。因此,所报道的相似性可能反映模型本身、测量过程,或两者兼有。我们对来自5个系列的17个模型(参数规模在0.6B至72B之间)的4种代表性神经科学启发范式进行了审计。我们的主要实验研究了概念方向的因果可操控性:使用原始激活单元、固定层和系数时,可操控性似乎随模型规模增加而提升,类似涌现能力;但该模式是由未校准的流程产生,而非操控文献中已确立的结论,其趋势取决于原始单元、读出指标和操作点三者的共同作用,修正其中任意一项都会消除该趋势。使用残差范数可比干预和预留操作点选择时,概念操控在所有规模下仍显著,但Qwen3系列无显著趋势,尽管置信区间未排除中等正斜率。其余结果参差不齐:线性地理世界图在所有测试的检查点(至72B)中均可稳定解码;数字量级被强编码,但单个神经元呈钟形或单调取决于选择准则;特定语言结构可定位,但跨语言不对称的方向在不同归因方法下会反转。这些结果表明,AI神经科学的主要限制并非缺乏现象,而是缺乏可比测量和充分控制。我们发布了方案、刺激和代码。
英文摘要
Large language models (LLMs) are increasingly reported to exhibit human-like neural and cognitive signatures, including concept cells, mental number lines, and cognitive maps. These claims often rely on linear probing and activation steering applied to a single model, yet both methods are highly sensitive to measurement choices. A reported parallel may therefore reflect the model, the measurement procedure, or both. We audit four representative neuroscience-inspired paradigms across 17 models from five families, spanning $0.6$B to $72$B parameters. Our main experiment examines the causal steerability of concept directions. With raw activation units and a fixed layer and coefficient, steerability appears to increase with model scale, resembling an emergent capability. However, this pattern is produced by an uncalibrated pipeline rather than by a claim established in the steering literature. The trend depends jointly on raw units, the readout metric, and the operating point; correcting any one of these removes it. With residual-norm-comparable interventions and held-out operating-point selection, concept steering remains significant at every scale, but shows no significant trend across the Qwen3 series, although the confidence interval does not rule out a moderate positive slope. The remaining results are mixed. A linear geographic world map is consistently decodable in every tested checkpoint up to $72$B. Number magnitude is strongly encoded, but whether individual neurons appear bell-shaped or monotonic depends on the selection criterion. Language-specific structure is localizable, but the direction of the cross-lingual asymmetry reverses under a different attribution method. These results suggest that the main constraint on AI neuroscience is not a lack of phenomena, but a lack of comparable measurements and adequate controls. We release the protocol, stimuli, and code.