arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

后归一化解码器Transformer中秩崩溃的机制诊断

Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure

Xingjian Wang, Qingyu Han, Xiaodong Luo, Yin Zhang

arXiv 2608.09417首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen; Shenzhen Research Institute of Big Data(香港中文大学(深圳); 深圳大数据研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究以token相似度为变量,分析后归一化解码器Transformer的秩崩溃机制,发现其前向相似度放大与反向修复无能的特性,实验验证了相关预测。

AI 中文摘要

仅解码器的深度Transformer常将原始后归一化(Post-Norm)架构替换为预归一化(Pre-Norm)变体,因为在常规初始化方案下,Post-Norm训练对预热(warmup)和学习率高度敏感。尽管已有研究将秩崩溃与梯度消失识别为相关现象,但人们仍不清楚因果注意力如何生成高相似度表示,以及训练动态为何无法修复这些问题。我们以token相似度作为标量状态变量,对Post-Norm秩崩溃开展两阶段分析:第一,初始化阶段,因果注意力近似充当前缀平均算子,会增加各层间的token相似度,而SwiGLU分支仅产生较小的阻尼效应;第二,一旦训练进入高相似度状态,预归一化残差范数的增长会使RMSNorm的反向传播因子具备收缩性,在温和条件下,到早期层的梯度会呈几何级数衰减。作为补充结果,我们刻画了崩溃网络的特性:其最优预测器是频率分布,且具有相对较高的损失下限,崩溃层的梯度在频率分布处消失。在C4数据集上训练的48层仅解码器Transformer实验,与预测的初始化阶段相似度增长及崩溃阶段梯度收缩相匹配,且显示崩溃运行始终接近预测的频率损失。这些结果共同区分了Post-Norm崩溃中的前向相似度放大与反向修复无能,同时刻画了崩溃网络的行为。

英文摘要

Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate under conventional initialization schemes. Although prior work has identified rank collapse and gradient vanishing as related symptoms, it remains poorly understood how causal attention creates high-similarity representations and why training dynamics fail to repair them. We give a two-stage analysis of Post-Norm rank collapse using token similarity as a scalar state variable. First, at initialization, causal attention acts approximately as a prefix-averaging operator that increases token similarity across depth, while the SwiGLU branch contributes only a smaller damping effect. Second, once training enters a high-similarity regime, growth of pre-normalization residual norms makes the RMSNorm backward factor contractive; under mild conditions, gradients to earlier layers decay geometrically. As a complementary result, we characterize the properties of a collapsed network: its best predictor is frequency distribution with relatively high loss floor, and gradients in collapsed layers vanish at frequency distribution. Experiments on 48-layer decoder-only Transformers trained on C4 dataset match the predicted initialization-time similarity growth and collapse-time gradient contraction, and show that collapsed runs stay near the predicted frequency loss. Together, these results distinguish the forward similarity amplification and backward repair incapacity in Post-Norm collapse, while also characterizing the behavior of collapsed networks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑