OR 否则:用于策略优化的可微信赖域
OR Else: A Differentiable Trust Region for Policy Optimization
- Quantiphi Inc(昆蒂菲公司)
- Self Machines Inc(自机器公司)
- The University of Texas at Arlington(德克萨斯大学阿灵顿分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究 PPO 和 GRPO 中裁剪代理目标导数突变问题,提出用输出重置(OR)规则优化。通过对比实验,在广义优势估计下 PPO-OR 有更高奖励模型得分,组相对优势下 GRPO-OR 虽平均得分未升但差异更小,OR 改变了优化行为但奖励效果有别。
AI中文摘要:
本文研究的近端策略优化算法(PPO)和广义信赖域策略优化算法(GRPO)基线使用裁剪代理目标,其有利方向饱和会导致标量目标导数出现突变。我们探讨输出重置(OR)这种平滑单边饱和规则是否能为大语言模型训练后提供有用的替代方案。PPO-OR 和 GRPO-OR 在展开相对令牌对数比率空间中用 OR 平方边际损失替换裁剪策略项;优势符号决定更新方向,令牌越过有利边际后直接 OR 残差为零。我们在广义优势估计(GAE)下将 PPO-clip 与 PPO-OR比较,在组相对优势下将 GRPO 与 GRPO-OR 比较,使用具有一个共享奖励模型且每种方法三个种子的 Anthropic hh-rlhf 上的 Llama-3.2-1B-Instruct。在 GAE 下,PPO-OR 的平均最终训练时奖励模型得分比 PPO-clip 高 0.305,种子间差异更大。在组相对优势下,GRPO-OR 平均得分不更高,但差异更小,终端 OR 残差接近零且过冲分数下降,而匹配的 GRPO 裁剪目标轨迹仍变化。两种组相对方法的展开到当前对数比率位移比 GAE 方法大得多,OR 不能持续减少它。因此,OR 在两种匹配比较中改变了优化行为,但观察到的奖励效果不同。在 G = 2 时,GRPO-OR 诊断结果未转化为奖励得分提升。更大的组是否会改变这一结果尚待研究。报告的分数是训练时奖励模型测量值,而非保留的人类偏好性能。
英文摘要:
PPO and the GRPO baseline studied here use clipped surrogate objectives whose favorable-direction saturation introduces an abrupt change in the scalar objective's derivative. We ask whether Output Reset (OR), a smooth one-sided saturation rule, offers a useful alternative for large language model post-training. PPO-OR and GRPO-OR replace the clipped policy term with an OR squared-margin loss in rollout-relative token log-ratio space; the advantage sign determines the update direction, and a token contributes zero direct OR residual after crossing the favorable margin. We compare PPO-clip with PPO-OR under generalized advantage estimation (GAE), and GRPO with GRPO-OR under group-relative advantages, using \texttt{Llama-3.2-1B-Instruct} on Anthropic \texttt{hh-rlhf} with one shared reward model and three seeds per method. Under GAE, PPO-OR has a mean final training-time reward-model score $0.305$ higher than PPO-clip, with a larger observed across-seed spread. Under group-relative advantages, GRPO-OR does not have a higher mean score, but shows a smaller observed spread, a near-zero terminal OR residual, and a declining overshoot fraction, while the matched GRPO clipped-objective trace remains variable. Both group-relative methods exhibit substantially larger rollout-to-current log-ratio displacement than the GAE methods, and OR does not consistently reduce it. Thus, OR changes optimization behavior in both matched comparisons, but the observed reward effect differs between them. At $G=2$, the GRPO-OR diagnostics do not translate into a reward-score gain. Whether larger groups change this outcome remains open. The reported scores are training-time reward-model measurements, not held-out human-preference performance.