arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16671cs.CL

LM 头是否会造成有害的梯度瓶颈?一项因果检验

Does the LM Head Create a Harmful Gradient Bottleneck? A Causal Test

Anand Murugan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过仅反向干预的因果检验发现,LM 头的几何压缩确实存在,但并非有害的优化瓶颈,且其效果显著优于同秩的因式分解前向头。

中文摘要 AI 辅助

语言模型(LM)头将宽度为 D 的隐藏状态映射到大小为 V 的词汇表,因此其转置最多可向 Transformer 返回 D 个独立方向。Godey 和 Artzi 认为,这种严重的投影是一种有害的优化瓶颈。我们将几何结构与因果主张分离开来:仅反向干预保留了普通 logits 和精确的 LM 头参数更新,仅降低发送到 Transformer 的梯度的秩。在字节级和 BPE-8192 WikiText-2 模型的五组配对随机种子实验中,降低反向秩会增加验证损失;而秩相同的因式分解前向头则会使损失大幅增加。在较大模型的半秩条件下,仅反向的损失增加为 0.0586(95% 置信区间 [0.0167, 0.1005]),而因式分解前向头的损失增加为 0.1795([0.1547, 0.2042])。词汇空间残差也对普通 LM 头的更新有贡献,移除该贡献是有害的。额外控制实验显示:重复标记故障与独立采样符号的数量混淆;添加从未作为目标的输出类别不会损害学习;投影诊断无法可靠预测实验进展;测试的辅助反馈路由未优于调优后的反向传播。这些结果证实存在强烈的几何压缩,但未证明其为有害的优化瓶颈。

英文摘要

The language-model head maps a hidden state of width D to a vocabulary of size V, so its transpose can return at most D independent directions to the Transformer. Godey and Artzi argue that this severe projection is a harmful optimization bottleneck. We separate the geometry from the causal claim. Our backward-only intervention keeps the ordinary logits and the exact LM-head parameter update while reducing only the rank of the gradient sent into the Transformer. Across five paired seeds on byte-level and BPE-8192 WikiText-2 models, reducing backward rank increases validation loss. An equally ranked factorized forward head, however, increases loss substantially more. At half rank in the larger model, the backward-only loss increase is 0.0586 (95% CI [0.0167, 0.1005]), while the factorized forward head increases loss by 0.1795 ([0.1547, 0.2042]). The vocabulary-space residual also contributes to the ordinary LM-head update, and removing that contribution is harmful. Additional controls show that repeated-token failures are confounded by the number of independently sampled symbols, that adding never-target output classes does not impair learning, and that projection diagnostics do not reliably predict progress in our runs. Tested auxiliary feedback routes do not beat tuned backpropagation. These results confirm strong geometric compression but do not establish that it is a harmful optimization bottleneck.

↑