发表机构
Renmin University of China(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型强化学习后训练中离策略数据问题,提出选择性重要性采样(SIS),受拒绝采样启发,以离策略模型为提议分布进行令牌级拒绝测试,理论证明可减少梯度估计差距,实验验证其有效性。
AI 中文摘要
大语言模型的强化学习后训练遵循“展开然后更新”的高效范式,不可避免地产生离策略训练数据。为解决此问题,提出重要性采样(IS),但令牌级比率在长序列上复合会导致严重的方差爆炸。我们提出选择性重要性采样(SIS),受拒绝采样启发。具体而言,SIS通过将离策略模型视为提议分布来实现,并进行令牌级拒绝测试:接受的令牌被视为策略内令牌,因此获得单位重要性分数,而被拒绝的令牌保留标准IS校正。我们提出的SIS在理论上被证明可以减少令牌级和序列级离策略梯度估计器之间的差距。SIS作为一个插件,只修改策略损失中的重要性比率,增加可忽略不计的挂钟开销,并且可以与各种RL后训练算法相结合。在数学和智能体基准测试中的密集和混合专家(MoE)大语言模型上的实验表明,SIS始终能改进所有目标,同时在离策略数据下提供更强的鲁棒性。
英文摘要
Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of "rollout then update", which inevitably results in off-policy training data. To resolve this, Importance sampling (IS) is proposed, while the token-level ratios compound over long sequences, causing severe variance exploded. A natural idea is "transferring" these off-policy token into on-policy token, so that the importance scores for correction are unnecessary. Following this idea, we propose Selective Importance Sampling (SIS), which is inspired by rejection sampling. Concretely, SIS implements by viewing off-policy model as proposal distribution, and implement a token-level rejection test: accepted tokens are viewed as on-policy, so that receive unit importance score, while rejected tokens retain the standard IS correction. Our proposed SIS is theoretically proved reducing the gap between token-level and sequence-level off-policy gradient estimators. The SIS acts as a plug-in that only modifies the importance ratio in the policy loss, adding negligible wall-clock overhead, and can be combine with a vast vary of RL post-training algorithms. Experiments on dense and MoE LLMs across math and agent benchmarks show that SIS consistently improves all objectives, while providing substantially stronger robustness under off-policy data.