发表机构
Peking University; National Engineering Research Center of Software Engineering, Peking University; School of Computer Science, Peking University; Key Laboratory of High Confidence Software Technologies, Ministry of Education; Center on Frontiers of Computing Studies, Peking University(北京大学; 北京大学软件工程国家工程研究中心; 北京大学计算机学院; 高可信软件技术教育部重点实验室; 北京大学计算前沿研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过NTK分析揭示稀疏奖励与密集蒸馏混合训练中的两种失败模式,提出M3系列方法结合幅度归一化与策略,以稳定在线策略蒸馏训练。
AI 中文摘要
基于可验证奖励的强化学习提供了一种稀疏的后训练信号:一个单一的二元结果评估整个轨迹,每个令牌接收到相同的序列级优势,无论其个体贡献如何。为了补充这种稀疏监督,越来越多的方法在策略梯度目标中添加一个标量加权的教师KL项,提供密集的令牌级指导,但在某些位置可能不可靠。尽管结合这些信号有好处,它们在优化过程中的交互可能会破坏联合训练的稳定性。为了理解这种不稳定性如何发展,我们通过神经正切核(NTK)分析研究混合奖励-蒸馏训练的学习动态。我们引入了交叉信号NTK $K_{DR}(n)$,这是一个令牌级统计量,用于衡量位置n处奖励梯度和蒸馏梯度之间的对齐程度。通过这一分析,我们识别出两种失败模式:1 幅度淹没,即奖励梯度超过蒸馏梯度数个数量级,因此即使弱的方向冲突也会导致蒸馏损失上升,尽管其明确包含在训练目标中;2 局部方向冲突,即序列级优势和教师的位置特定分布在同一令牌处引发相反的更新($K_{DR}(n)\\!<\\!0$)。这些效应的严重程度取决于优化机制:梯度范数比 $\kappa\\!=\\!\\|\nabla\mathcal{L}_R\\|/\\|\nabla\mathcal{L}_D\\|$ 在不同任务中大约变化一个数量级,我们的实验揭示了一个经验阈值,超过该阈值,朴素混合可能导致持续的训练崩溃。受这些发现的启发,我们引入了M3系列,它结合了幅度归一化与三种策略...
英文摘要
Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a scalar-weighted teacher KL term to the policy-gradient objective, providing dense token-level guidance that may be unreliable at some positions. Despite the benefits of combining these signals, their interaction during optimization can destabilize joint training. To understand how this instability develops, we study the learning dynamics of hybrid reward--distillation training through a neural tangent kernel (NTK) analysis. We introduce the cross-signal NTK $K_{DR}(n)$, a token-level statistic that measures the alignment between reward and distillation gradients at position n. Through this analysis, we identify two failure modes: 1 Magnitude drowning, where the reward gradient exceeds the distillation gradient by orders of magnitude, so that even weak directional conflict can cause the distillation loss to rise despite its explicit inclusion in the training objective; and 2 Localized directional conflict, where the sequence-level advantage and the teacher's position-specific distribution induce opposing updates at the same token ($K_{DR}(n)\!<\!0$). The severity of these effects depends on the optimization regime: the gradient-norm ratio $κ\!=\!\|\nabla\mathcal{L}_R\|/\|\nabla\mathcal{L}_D\|$ varies by roughly an order of magnitude across tasks, and our experiments reveal an empirical threshold beyond which naive mixing can lead to persistent training collapse. Motivated by these findings, we introduce the M3 family, which combines magnitude normalization with three strategies...