发表机构
Institute of Foundation Models; USC; CMU(基础模型研究院; 南加州大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出学习深度先验与正交注入,优化循环语言模型的不动点,实现更快的预填充、解码和RL训练,并在各规模下降低困惑度。
AI 中文摘要
循环语言模型的每一次循环都会在训练、解码、预填充和强化学习(RL)中增加成本。循环状态越接近不动点,到达不动点的路径就越不重要。这使得训练中可以采用截断反向传播;终端键值(KV)共享用于解码,且几乎不损失准确性;蒸馏学生模型的预填充速度最高可提升1.79倍;以及从保存的轨迹状态计算梯度的RL更新,比通过重放轨迹进行反向传播快2倍。因此,我们改进了塑造这些不动点的两个训练组件:深度先验和输入注入。固定深度训练破坏了KV共享,而Huginn的宽泛深度先验支持共享,但在目标深度上对监督的稀释超过了共享所需;我们从预测反馈中学习先验,并引入熵项以保持其宽泛性。现有的注入方案允许状态沿输入方向的分量放大或抵消注入;我们通过正交注入移除该分量。从100M到1.6B参数,相对于Huginn的先验和现有注入方案,学习到的先验和正交注入在每个规模上都降低了困惑度。在1.6B规模下,使用3倍更小的KV缓存的学习先验在下游平均性能上与使用完整缓存的固定深度训练相匹配。
英文摘要
Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss in accuracy; a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout states, 2x faster than backpropagating through the replayed trajectory. We therefore improve the two components of training that shape these fixed points: the depth prior and input injection. Fixed-depth training breaks KV sharing, and Huginn's broad depth prior supports sharing but dilutes supervision at the target depth more than sharing requires; we learn the prior from prediction feedback, with an entropy term that keeps it broad. Existing injection schemes let the state's component along the input amplify or cancel the injection; we remove this component with orthogonal injection. From 100M to 1.6B parameters, the learned prior and orthogonal injection lower perplexity at every scale relative to Huginn's prior and existing injection schemes, respectively. At 1.6B, the learned prior with a 3x smaller KV cache matches the downstream average of fixed-depth training with the full cache.
CommentsCode and checkpoints: https://github.com/ifm-ai/xllm-loop