重新思考大语言模型的归一化位置:课程深度增长下的后归一化
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing
AI总结:
本研究在Qwen3-8B到9层模型的蒸馏任务中发现,课程深度增长下后归一化优于预归一化,二者表现存在数量级差异,证明归一化位置与训练课程是耦合设计选择。
AI中文摘要:
预归一化是现代Transformer中的标准归一化位置,因为它便于全深度模型的联合优化。本文探究当通过课程引入深度时,这种偏好是否仍然存在。在课程深度增长中,每个新增的Transformer块都会接收由已训练前缀生成的边界表示,这使得归一化位置与前向条件化相关。因此,本文测试了归一化位置与训练课程是否存在交互。在以Qwen3-8B为教师模型、9层模型为学生模型的受控蒸馏研究中,联合训练下预归一化与后归一化的验证交叉熵仅相差0.0004,表现无显著差异;而在课程深度增长设置下,后归一化相比预归一化提升了0.0328,差距达到一个数量级。由学生模型活跃层令牌匹配的后联合控制组表现仍差于后增长组,这排除了计算量作为唯一解释。在课程过程中,两种归一化位置的表现排名发生交叉:新增块后,后归一化开始占据优势。单块控制和冻结控制将排名变化定位到块新增环节,而非浅层块质量或重新训练。边界诊断显示后归一化与稳定的残差尺度相关,预归一化则与结构令牌尺度漂移相关;在固定批次下,新增块前的最终块几乎是恒等映射。结合阶段式交叉的结果,这些观察与新增块后的边界尺度条件化一致。本研究的结果促使在该蒸馏设置中将归一化位置和训练课程视为耦合的设计选择。
英文摘要:
Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning. We therefore test whether placement and training curriculum interact. In a controlled distillation study with a Qwen3-8B teacher and a nine-layer student, pre-norm and post-norm are indistinguishable under joint training, differing by $0.0004$ validation CE, while post-norm improves over pre-norm by $0.0328$ under curriculum growth, an order of magnitude larger. A post-joint control matched by student active-layer tokens remains worse than post-grow, which rules out compute as the sole explanation. The ranking crosses over during the curriculum: post-norm takes the lead once blocks are appended. Single-block and freeze controls localize the ranking change to block appending rather than shallow-block quality or retraining. Boundary diagnostics associate post-norm with stable residual scales and pre-norm with structural-token scale drift; on a fixed batch, the final pre-grow block is also nearly identity-mapped. Together with the phase-wise crossover, these observations are consistent with boundary-scale conditioning after new blocks are appended. The results motivate treating normalization placement and training curriculum as coupled design choices in this distillation setting.