发表机构
University of Rochester(罗切斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LayerRoPE通过将残差流上的逐层归一化权重γ替换为共享向量与深度条件标量,将深度增长转化为隐式深度位置编码,在减少参数和FLOPs的同时,以3.4倍更少计算达到Pre-Norm的1.3B损失,并提升学习率敏感性3-10倍。
AI 中文摘要
当数据在Transformer中传播时,其隐藏状态的范数随深度呈数量级增长,这一现象被称为“深度诅咒”,且几乎普遍被视为需要抑制的病态。我们持相反观点。在来自9个家族的16个预训练大语言模型中,涵盖密集、混合专家和混合架构以及Pre-Norm、Peri-Norm和Post-Norm设计,我们发现这种增长反映了一种涌现的深度位置编码,由残差流上唯一学习的逐层增益——归一化权重γ承载:随深度增加,γ在幅度上增长并在方向上旋转,共同编码层索引。我们通过LayerRoPE使这种深度条件编码显式化,它是沿深度轴的RoPE的隐式类比,将所有逐层γ向量替换为单个共享向量和深度条件标量,净减少参数且FLOPs变化小于0.02%。在扩展到100B+令牌的模型阶梯上,LayerRoPE持续优于Pre-Norm、Post-Norm、Peri-Norm和Layer-Norm Scaling,达到Pre-Norm的1.3B损失所需计算量减少3.4倍;LayerRoPE是唯一在深度扩展到512层时表现出强收敛且近乎单调改进的方法。它将学习率敏感性提高了3-10倍,并可直接迁移至循环潜变量模型和视觉Transformer,且持续改进它们。检查其学习到的调度推翻了主流前提:LayerRoPE不缩小残差流,而是加宽它,抑制每个块读取的内容,同时放大其写入的内容。我们的结果表明,深度稳定性需要的不是抑制残差流,而是对其所馈送的计算块进行深度条件调节。
英文摘要
As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed. We take the opposite view. Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight $γ$: with depth, $γ$ grows in magnitude and rotates in direction, jointly encoding the layer index. We make this depth-conditioned encoding explicit with LayerRoPE, an implicit analog of RoPE along the depth axis, which replaces all layerwise $γ$ vectors with a single shared vector and depth-conditioned scalars, at a net reduction in parameters and $<0.02\%$ change in FLOPs. Across a model ladder scaled up to $100$B+ tokens, LayerRoPE consistently outperforms Pre-, Post- and Peri-Norm and Layer-Norm Scaling, reaching Pre-Norm's 1.3B loss with $3.4\times$ less compute; LayerRoPE is the only approach that shows strong convergence and improves near monotonically as depth scales to 512 layers. It improves learning-rate sensitivity by $3$-$10\times$, and transfers naively to and consistently improves looped latent models and Vision Transformers. Inspecting its learned schedule inverts the prevailing premise: LayerRoPE does not shrink the residual stream but widens it, damping what each block reads while amplifying what it writes. Depth stability, our results suggest, calls not for suppressing the residual stream, but for depth-conditioned regulation of the computational blocks it feeds.