发表机构
INSAIT, Sofia University “St. Kliment Ohridski”; National Research Center for Applied Cybersecurity ATHENE(INSAIT,圣克莱门特奥赫里德斯基索非亚大学; ATHENE国家应用网络安全研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对LLM强化学习的探索问题,提出参数空间探索思路,引入3PO方法,在OLMo-3-1025-7B等模型的数学推理等任务上,以相近计算成本相比GRPO提升性能,减少训练中的异常分组与轨迹。
AI 中文摘要
长期以来,探索一直是强化学习研究的重点。近来越来越多的证据表明,探索也是大语言模型(LLM)强化学习方案中的重要组成部分,可显著影响下游任务性能。现有诸多方法在动作空间层面控制探索,例如采用温度缩放,但这类方法无法对token进行重新排序,仅能影响输出分布的方差,这限制了探索效果,可能导致训练发散或停滞。本文研究参数空间探索,其中通过从后验分布中采样不同策略来生成轨迹,各策略可探索不同轨迹;采样更多或更少多样化的策略是对探索的补充控制手段。我们提出一类名为扰动参数策略优化(3PO)的方法,该方法采用不同采样策略及不同轨迹分组方式进行奖励估计。在OLMo-3-1025-7B和Qwen2.5-Math-7B上针对数学推理与代码生成任务开展的实验显示,这些方法在浮点运算量(FLOPs)成本几乎相同的情况下,相比标准GRPO可持续提升下游平均性能;此外,在训练过程中,使用多个参数采样相比GRPO及动作空间基线方法,会产生更少的零优势分组、畸形或错误的轨迹。总体而言,本研究证明参数空间探索可改进大语言模型的强化学习效果。
英文摘要
Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performance. Many existing methods control exploration in the action-space, for example, using temperature scaling. However, these methods cannot reorder tokens but only influence the variance in the output distribution. This limits exploration and can lead to divergence or stalled training. Here, we investigate parameter-space exploration, where rollouts are generated by sampling different policies from a posterior that may each explore different rollouts. Sampling less or more diverse policies is then a complementary control lever over exploration. We introduce a family of methods called Perturbed Parameter Policy Optimization (3PO) which use different sampling strategies and different rollout grouping for reward estimation. Experiments on OLMo-3-1025-7B and Qwen2.5-Math-7B across mathematical reasoning and code generation tasks show that these approaches consistently improve average downstream performance over standard GRPO at a near-identical FLOPs cost. Moreover, using multiple parameter samples consistently produces fewer zero-advantage groups and malformed or incorrect rollouts during training than GRPO and action-space baselines. Overall, our work presents evidence that parameter-space exploration can improve reinforcement learning for LLMs.