arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18612cs.LGcs.CL

弱化神经元:Transformer 中输入-输出功能性的超常影响

Weakening Neurons: An Input-Output Functionality in Transformers with Outsize Influence

Sebastian Gerstner, Hilal AlQuabeh, Kentaro Inui, Hinrich Schütze

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出基于余弦相似度的分析方法,发现LLM中弱化神经元虽少但频繁激活且影响大,尤其在负门控值时显著,揭示其独特功能。

中文摘要 AI 辅助

我们分析了大型语言模型(LLMs)中基于GLU的神经元的学习输入-输出行为。我们提出了一种简单的分析方法:对于每个神经元,我们计算其输入(读取)和输出(写入)权重向量之间的余弦相似度。在这种方案中,强负余弦相似度表明该神经元弱化了其在残差流中检测到的方向,因此我们将其称为弱化神经元。这使我们获得了许多新的见解。首先,我们展示了九个不同的LLM具有相似的模式:弱化神经元主要出现在后期层,而其对应物(条件性)强化神经元则频繁出现在中早期层。其次,我们发现弱化神经元表现出令人惊讶的行为:尽管数量很少,但它们经常被激活,并对模型行为产生很大影响。第三,当门控值为负时,弱化神经元对模型输出有强烈影响——这令人惊讶,因为负门控值通常不被预期编码功能。

英文摘要

We analyze the learned input-output behavior of GLU-based neurons in large language models (LLMs). We propose a simple analysis method: For each neuron, we compute the cosine similarities between its input (reading) and output (writing) weight vectors. In this scheme, a strong negative cosine similarity indicates the neuron weakens the direction it detects in the residual stream, so we call this a weakening neuron. This allows us to gain a number of novel insights. First, we show that nine different LLMs have similar patterns: weakening neurons appear mostly in late layers whereas their counterparts, (conditional) strengthening neurons, are frequent in early-middle layers. Second, we find that weakening neurons display surprising behavior: even though there are few, they activate often and have a large influence on model behavior. Third, weakening neurons have a strong effect on model output when gate values are negative -- which is surprising since negative gate values are not expected to encode functionality.

发表机构

  • LMU Munich(慕尼黑大学)
  • Munich Center for Machine Learning(慕尼黑机器学习中心)
  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)
  • Tohoku University(东北大学)
  • RIKEN(理化学研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑