arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20807cs.LG

分数中心化稳定离策略强化学习

Score Centering Stabilizes Off-policy Reinforcement Learning

Martin Marek, Max Ryabinin

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对大型语言模型强化学习中的训练-推理不匹配问题,提出加性分数中心化校正项以消除漂移,稳定训练,并在0.6B至30B参数规模下匹配或超越重要性采样方法。

中文摘要 AI 辅助

大型语言模型的强化学习(RL)对训练引擎与推理引擎之间的微小差异极为敏感,这种差异通常被称为训练-推理不匹配(TIM)。然而,完全消除TIM是不切实际的,因为这会以牺牲推演效率为重大代价。在本文中,我们表明TIM下强化学习的不稳定性主要由漂移引起:训练与推理引擎之间持续存在的偏差,且随着每一步训练而累积。我们推导出一个加性的“分数中心化”校正项,通过消除漂移来稳定TIM下的强化学习。当训练从0.6B到30B参数的模型时,仅使用分数中心化就能在量化条件下匹配或超越基于重要性采样的方法,且随着不匹配程度加剧,差距逐渐扩大。由于该校正项是加性的,分数中心化还能与重要性采样组合使用——在我们的陈旧性实验中,它们的组合优于纯重要性采样基线。

英文摘要

Reinforcement learning (RL) of large language models is notoriously sensitive to small differences between training and inference engines, often referred to as the training-inference mismatch (TIM). However, completely eliminating TIM is impractical, as it would come at a major cost to rollout efficiency. In this paper, we show that the instability of RL under TIM is primarily caused by drift: a persistent bias between training and inference engines that accumulates with every training step. We derive an additive "score centering" correction term that stabilizes RL under TIM by canceling drift. When training models from 0.6B to 30B parameters, score centering alone matches or outperforms methods based on importance sampling under quantization, with the gap growing as the mismatch becomes more severe. Because the correction is additive, score centering also composes with importance sampling -- their composition outperforms pure importance-sampling baselines in our staleness experiments.

发表机构

  • Together AI

机构由 AI 辅助整理,请以论文原文为准。

↑