Z-Loss在稠密输出头和稀疏路由器中的反向几何
Z-Loss Backward Geometry in Dense Output Heads and Sparse Routers
浏览论文内容
中文总结 AI 辅助
本文从反向传播视角分析Z-loss,提出反向传输诊断框架,分离源幅度与传输因子,解释梯度差异,并在GPT-2和Pythia模型上验证架构感知变体可减少尾部并保持困惑度。
中文摘要 AI 辅助
Z-loss已被广泛应用于语言模型输出头和稀疏混合专家路由器的logits上。Z-loss约束这些输出头和路由器的softmax对数归一化因子,从而限制大logit偏移,减少有限精度舍入暴露,并避免训练损失发散。这些用例出现在现代Transformer设置中,其中大词汇量softmax头、top-$k$路由、融合损失和混合精度优化器相互作用。Z-loss通常仅被理解为对对数归一化因子的标量惩罚。本文则从反向传播的角度分析Z-loss,重点关注Z-loss惩罚产生的梯度。logit空间梯度(我们称之为反向源)被注入到Z-loss分支反向传播的logit边界;因此,反向源的效果取决于梯度传输所通过的架构和实现。我们为Z-loss开发了一种反向传输视图,将源的标量幅度和softmax形状与传输因子分离。这些因子包括公共偏移坐标、绑定嵌入路径、输出到隐藏增益、融合损失源一致性、面向优化器的更新以及top-$k$路由器缩减规模。这些诊断表明,几乎相同的前向Z-loss值可以与不同的logit空间Z-loss梯度共存,并且在架构和优化器传输后,产生不同的参数更新。传输诊断还解释了为什么原始logit Z-loss可以减少标量尾部而不改变输出到隐藏增益,以及为什么活跃路由缩减会改变有效路由器系数。在GPT-2和Pythia系列模型对WikiText-103和FineWeb-Edu的评估中,架构感知变体减少了反向几何尾部,同时在低系数机制中保持了可比较的验证困惑度。
英文摘要
Z-loss has been widely applied to the logits of language-model output heads and sparse mixture-of-experts routers. Z-loss constrains the softmax log-normalizers of these output heads and routers, thereby limiting large-logit excursions, reducing finite-precision roundoff exposure, and avoiding training-loss divergence. These use cases arise in modern Transformer settings where large-vocabulary softmax heads, top-$k$ routing, fused losses, and mixed-precision optimizers interact. Z-loss has typically been understood only as a scalar penalty on the log-normalizer. This paper instead analyzes Z-loss from a backward-pass perspective, focusing on the gradients produced by the Z-loss penalty. The logit-space gradient, which we call the backward source, is injected at the logit boundary of the Z-loss branch of backpropagation; consequently, the backward source's effect depends on the architecture and implementation through which the gradient is transported. We develop a backward-transport view for Z-loss that separates the source's scalar amplitude and softmax shape from the transport factors. These factors include common-shift coordinates, tied-embedding pathways, output-to-hidden gain, fused-loss source consistency, optimizer-facing updates, and top-$k$ router reduction scale. These diagnostics show that nearly identical forward Z-loss values can coexist with distinct logit-space Z-loss gradients and, after architectural and optimizer transport, distinct parameter updates. The transport diagnostics also explain why raw-logit Z-loss can reduce scalar tails without changing output-to-hidden gain and why active-route reductions alter the effective router coefficient. Across evaluations of models in the GPT-2 and Pythia families on WikiText-103 and FineWeb-Edu, architecture-aware variants reduce backward-geometry tails while maintaining comparable validation perplexity in low-coefficient regimes.