发表机构
Tencent Hunyuan; UIUC; NUS(腾讯混元; 伊利诺伊大学厄巴纳 - 香槟分校; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型强化学习中PPO方向标准不足,提出预测性散度掩码,通过判断策略梯度步骤对散度的影响改进训练,针对离散softmax策略推导预测并开发轻量级估计器,提升不同模型规模和精度设置下的强化学习训练效果。
AI 中文摘要
大语言模型的强化学习通常依靠信任区域掩码来稳定离策略更新。主流的近端策略优化(PPO)方法使用采样令牌重要性比率来满足两个标准:一个是接近度标准,即判断策略是否偏离行为策略太远;另一个是方向标准,即判断更新是否使其离行为策略更远)。近期的DPPO工作改进了接近度标准,但方向标准仍继承自PPO。我们观察到基于比率的方向标准是一个单样本代理,可能与定义接近度标准的散度变化符号不一致。因此,我们提出了预测性散度掩码,它判断下一个策略梯度步骤是否会增加或减少信任区域使用的相同散度。对于大语言模型强化学习中使用的离散softmax策略,我们以封闭形式得出此预测。由于生产部署引擎仅暴露词汇表的截断(top-K)视图,我们为此预测开发了两个轻量级的top-K估计器。详细分析表明,基于散度的方向比采样比率更符合散度的实际变化,由此产生的掩码在不同模型规模和精度设置下都能改善强化学习训练。
英文摘要
Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize off-policy updates. The dominant PPO-style approach uses the sampled-token importance ratio for two criteria: a proximity criterion, which asks whether the policy has moved too far from the behavior policy, and a direction criterion, which asks whether the update pushes it farther away. Recent work DPPO improves the proximity criterion by replacing PPO's ratio-based test with a probability divergence between the behavior and training policies. However, its direction criterion is still inherited from PPO. A token can be masked only when the sampled-token importance ratio moves away from one. We observe that this ratio-based direction criterion is a single-sample proxy that can disagree in sign with the change of the divergence that defines the proximity criterion. We therefore propose the predictive divergence mask, which asks whether the next policy-gradient step will increase or decrease the same divergence used by the trust region. For the discrete softmax policies used in LLM RL, we derive this prediction in closed form. Because production rollout engines expose only a truncated (top-K) view of the vocabulary, we develop two lightweight top-$K$ estimators for this prediction. Detailed analysis shows the divergence-based direction is better aligned with the realized change of the divergence than the sampled ratio, and the resulting masks improve RL training across model scales and precision settings.