发表机构
Tohoku University; SOKENDAI; NINJAL; MBZUAI; RIKEN; Preferred Networks, Inc.(东北大学; 综合研究大学院大学; 日本国立国语研究所; Mohamed bin Zayed 人工智能大学; 日本理化学研究所; Preferred Networks公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示注意力汇聚与大规模激活源于因果掩码引发的自集中和值非混合,而非RoPE,为量化策略提供新见解。
AI 中文摘要
大型语言模型(LLMs)在序列的初始位置常常表现出“注意力汇聚”(Attention Sink,AS)以及伴随的“大规模激活”(Massive Activations,MAs)。这些现象经常同时出现,且MAs可能对低比特量化构成挑战。在本研究中,我们分析了无论占据初始位置的词元是什么,AS和MAs在该位置出现的潜在因素。我们的实验表明,由因果掩码导致的注意力自集中,以及随后注意力输出中的值非混合,共同促成了AS和MAs。这些发现为LLMs的内部动态提供了新的经验证据,为未来的量化策略提供了见解,并增进了我们对注意力层内部机制的理解。
英文摘要
Large Language Models (LLMs) often exhibit "Attention Sink" (AS) and the accompanying "Massive Activations" (MAs) at the initial position of a sequence. These phenomena frequently co-occur, and MAs can pose challenges for low-bit quantization. In this study, we analyze the factors underlying AS and MAs that emerge at the initial position regardless of the token occupying it. Our experiments suggest that self-concentration of attention, resulting from the causal mask, and the subsequent Value-non-mixing in attention outputs contribute to AS and MAs. These findings provide new empirical evidence on the internal dynamics of LLMs, offering insights that may inform future quantization strategies and advance our understanding of the internal mechanisms of attention layers.
CommentsAccepted at EMNLP 2026