发表机构
Startlux; Tsinghua University; University of Chinese Academy of Sciences; The University of Hong Kong; University of Sydney; Columbia University(星创公司; 清华大学; 中国科学院大学; 香港大学; 悉尼大学; 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统揭示混合线性注意力大语言模型中大规模激活的两种形态,证实其在多架构、多配置等场景下的重复性,明确其对输出门控的响应规律及机制,为相关模型优化提供依据。
AI 中文摘要
我们首次对层间混合线性注意力(HLA)大语言模型中的大规模激活(MAs)开展系统性研究,揭示了两种与架构对齐的形态:大规模激活会在全注意力层前立即形成注意力前峰值(PAS),并可在中间的线性注意力层中持续存在,从而产生峰间平台期(ISP)。随着全注意力变得更密集,连续的注意力前峰值会通过峰间平台期日益相连,最终恢复全注意力大语言模型的稳定大规模激活形态。我们在5种线性注意力架构、6种混合配置、5种数据领域及总参数规模12亿至3970亿的代表性开源混合模型中,证实了该组织模式的重复性。对基于GDN的混合模型开展的规模达13亿的受控预训练显示,两种形态均较早出现,且对输出门控的响应不对称:全注意力输出门控会大幅削弱其绝对幅值,但不会消除其层级组织;移除GDN门控则会产生相对温和的放大。从机制上看,我们的系统性异常值分析支持由大规模激活的取消时机所决定的共享生命周期解释:注意力前峰值遵循局部写入-汇聚-取消过程,而峰间平台期的持续存在与延迟取消一致;在全注意力极限下,该解释可恢复全注意力大语言模型的稳定大规模激活形态。我们的代码可在该https URL获取。
英文摘要
We present the first systematic study of massive activations (MAs) in layer-interleaved Hybrid linear attention large language models (HLA LLMs), examining their architectural organization, training-time emergence, underlying mechanisms, and functional significance. Across five linear attention architectures, six hybridization configurations, and five input domains, we identify two architecture-aligned morphologies: pre-attention spikes (PAS) immediately before full attention and inter-spike plateaus (ISP) persisting through intervening linear attention layers. Denser full attention increasingly connects PAS through ISP, approaching the persistent MAs of conventional Transformers. This organization also recurs across 12 public checkpoints spanning 1.2B-397B parameters, covering linear attention and state-space hybrids. Controlled pretraining of Gated DeltaNet (GDN) hybrids up to 1.3B reveals early emergence and consolidation of both morphologies, alongside asymmetric gating effects. Specifically, full attention output gates strongly attenuate MA magnitudes without eliminating their organization, whereas removing GDN output gates yields modest amplification. Mechanistically, we develop a shared systematic-outlier account: PAS follows a localized write-sink-cancel process, while ISP is consistent with delayed cancellation. Functionally, our interventions show that deleting only the four largest-magnitude PAS coordinates at each full attention input reduces mean downstream accuracy by 21.9%-63.6% relative to normal inference. Moreover, reference-conditioned spike-to-plateau connection consistently improves mean real-world retrieval accuracy, yielding relative gains of 1.1%-12.6% without retraining. Our code is available at https://github.com/StartLuxLabs/Massive-Activations-HLA.
CommentsUnder review