arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EasyPPO:稳定评论家是关键

EasyPPO: Stabilizing the Critic Is Key

Xuanyi Zhou, Qiuyang Mang, Huanzhi Mao, Dacheng Li, Wenhao Chai, Mayank Mishra, Yichuan Wang, Karthik Narasimhan, Alvin Cheung, Joseph E. Gonzalez

arXiv 2609.36802首次发表:更新:

发表机构

University of California, Berkeley; Princeton University(加州大学伯克利分校; 普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EasyPPO通过仅过滤演员的过长轨迹、噪声归一化评论家回归和更小的评论家小批量,稳定了PPO中的评论家,在多个任务上显著优于基线。

AI 中文摘要

近端策略优化(PPO)的一个关键优势是其学习的评论家,它利用强化学习过程中收集的历史轨迹来估计期望回报并减少策略梯度方差。然而,我们发现评论家也是大型语言模型(LLMs)强化学习中不稳定的主要来源。我们识别出两种使PPO不稳定的评论家故障模式。首先,从演员和评论家中过滤截断的轨迹会使策略目标偏向于基于完成度的奖励,导致即使条件奖励改善,截断率也可能增加。其次,异质的回报噪声可能导致高方差提示在有限批次中主导评论家更新。我们引入EasyPPO来解决这些故障。仅演员的过长过滤训练评论家使用来自已完成和截断轨迹的回报。噪声归一化的评论家回归通过每个提示的采样回报的标准差的倒数来加权该提示的评论家损失,从而平衡各提示间的噪声贡献。适度更小的评论家小批量在梯度裁剪期间将异常值的影响限制在更少的轨迹上。在FrontierCS上的连续奖励编码、AIME24上的二元奖励数学推理以及Search-R1上的多轮搜索中,EasyPPO在整个训练周期内保持稳定,并且始终优于vanilla PPO、VAPO和HL-Gauss PPO。其最佳验证分数相对于PPO分别相对提升了14.89%、2.28%和9.47%。

英文摘要

A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for large language models (LLMs). We identify two critic failure modes that destabilize PPO. First, filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, allowing truncation to increase even as conditional reward improves. Second, heterogeneous return noise can cause high-variance prompts to dominate critic updates in finite batches. We introduce EasyPPO to address these failures. Actor-only overlong filtering trains the critic on returns from both completed and truncated rollouts. Noise-normalized critic regression weights each prompt's critic loss by the inverse standard deviation of its sampled returns, balancing noise contributions across prompts. Moderately smaller critic mini-batches confine outlier influence to fewer rollouts during gradient clipping. Across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24, and multi-turn search on Search-R1, EasyPPO remains stable throughout the full training horizon and consistently outperforms vanilla PPO, VAPO, and HL-Gauss PPO. Its best validation scores show relative gains of 14.89%, 2.28%, and 9.47% over PPO, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑