发表机构
Indian Institute of Technology Bombay; Jahangirnagar University; National Institute of Technology, Silchar(印度孟买理工学院; 贾汉吉尔纳加尔大学; 锡尔恰尔国立理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型中算术启发式神经元在符号算术、自然语言文字问题和Python代码中是否形式不变。通过两阶段管道识别神经元,发现跨格式有共享电路,其对后期算术计算必要且充分,跨格式失败因激活状态,共享神经元所属启发式家族一致。
AI 中文摘要
大语言模型在问题的一种表述上成功,却在等效表述上失败,其原因未知。近期研究表明算术由‘启发式集合’产生,由代表不同算术策略的稀疏MLP神经元编码。本文在三个Llama - 3模型中研究算术启发式神经元在符号算术、自然语言文字问题和Python代码中是否形式不变。通过两阶段管道识别神经元,发现所有格式共享一组紧凑神经元,针对性干预表明该共享电路对后期算术计算既必要又充分。跨格式失败源于激活状态而非不同电路,且共享神经元跨格式一致属于相同启发式家族,证明大语言模型中算术计算在神经元层面很大程度上形式不变。
英文摘要
Large language models often succeed on one formulation of a problem while failing on an equivalent formulation. Whether these failures arise from distinct internal circuits or different activation states of a shared circuit remains unknown. Recent mechanistic interpretability studies suggest that arithmetic in LLMs emerges from a "bag of heuristics," encoded by a sparse set of MLP neurons that represent distinct arithmetic strategies. We investigate whether arithmetic heuristic neurons are form-invariant across symbolic arithmetic, natural language word problems, and Python code in three Llama-3 models. In each format, we identify arithmetic heuristic neurons using a two-stage pipeline combining attribution patching and activation patching. A compact set of neurons is shared across all three formats, and targeted interventions show this shared circuit is both necessary and sufficient for late-layer arithmetic computation. Transferring the shared neurons' activations from a successful execution in one format to a failed execution in another recovers most incorrect predictions, exceeding 97% for addition and subtraction, indicating that cross-format failures arise from activation states rather than distinct circuits. Moreover, shared neurons consistently belong to the same heuristic families across formats, demonstrating that arithmetic computation in LLMs is largely form-invariant at the neuron level.
CommentsUnder Review