发表机构
University of Delaware; Peking University(特拉华大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发现大规模激活由输入嵌入中的单一门控通道(MAGC)控制,并提供了理论解释,验证了其在多个LLM中的存在与效应。
AI 中文摘要
大规模激活现象,即少数隐藏通道表现出异常大的幅度,在大语言模型(LLMs)中普遍存在。然而,一个token在通过预训练的LLM传播时如何产生大规模激活的机制仍知之甚少。本文发现,大规模激活的出现由输入嵌入到尖峰前馈网络(FFN)中的一个单一通道控制。该通道的位置对于特定的LLM是固定的。我们将此通道命名为大规模激活门控通道(MAGC)。当MAGC的值足够大(或足够小,取决于LLM)时,尖峰FFN的输出表现出大规模激活。通过检查四个模型家族和不同模型大小的六个LLM,我们验证了MAGC的存在和效应。我们进一步提供了MAGC诱导大规模激活机制的理论解释。当MAGC的值足够大(或足够小)时,尖峰FFN的输出渐近地简化为二次形式,该形式混合了FFN下投影矩阵的几列。由于这些列表现出大规模激活的形状,因此输出表现出大规模激活。
英文摘要
Massive activations, a phenomenon in which a small number of hidden channels exhibit exceptionally large magnitudes, are pervasive in large language models (LLMs). However, the mechanism by which a token develops massive activations as it propagates through a pretrained LLM remains poorly understood. In this paper, we find that the emergence of massive activations is controlled by a single channel in the input embedding to a spike feed-forward network (FFN). The position of this channel is fixed for a particular LLM. We name this channel the massive activation gating channel (MAGC). When the value of the MAGC is sufficiently large (or small, depending on the LLM), the output of the spike FFN exhibits massive activations. Examining six LLMs across four model families and different model sizes, we verify the existence and effect of MAGC. We further provide a theoretical explanation of the mechanism by which MAGC induces massive activations. When the value of MAGC is sufficiently large (or small), the output of a spike FFN asymptotically reduces to a quadratic form that mixes a few columns of the down-projection matrix of the FFN. Since these columns exhibit the shape of massive activations, the output therefore exhibits massive activations.