重新思考PPO中的评论家学习:理解并缓解价值扁平化
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
- Shanghai Jiao Tong University(上海交通大学)
- Shanghai AI Laboratory(上海人工智能实验室)
- Westlake University(西湖大学)
- Nanjing University(南京大学)
- Tsinghua University(清华大学)
- The Chinese University of Hong Kong(香港中文大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文发现PPO评论家存在价值扁平化问题,提出SP$^3$O算法仅监督少数分离状态,在Qwen3-Base上缓解该问题并提升策略性能。
AI中文摘要:
在大型语言模型的强化学习中,近端策略优化(PPO)通常使用评论家(critic)来估计状态价值并减少策略更新的方差。然而,我们发现PPO评论家存在一个系统性的失败模式,我们称之为“价值扁平化”(Value Flattening):从多个蒙特卡洛延续中估计的状态价值在中间状态间急剧变化,而评论家的预测却相对平坦。我们进一步在受控的FrozenLake环境中观察到这一现象,并发现随着状态空间的增大,该现象变得更加明显。我们的理论和实证分析将价值扁平化与评论家损失中的隐式方差惩罚以及来自时间相关且梯度相似状态的冗余更新联系起来。基于这些发现,我们提出了稀疏近端策略优化(SP$^3$O),该算法仅对每个响应中的少数几个分离良好的状态应用价值损失,以减轻这两种效应。在Qwen3-Base上的实验表明,SP$^3$O仅需在每个响应中监督三个状态,即可缓解价值扁平化,并在不同模型规模和评估套件中持续改进学习到的策略。综合来看,我们的结果将价值扁平化识别为标准PPO中评论家学习的一个重要但被忽视的失败模式,并表明一种简单的稀疏监督策略可以缓解该问题。
英文摘要:
In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Carlo continuations, change sharply across intermediate states while critic predictions remain comparatively flat. We further observe this phenomenon in a controlled FrozenLake environment and find that it becomes more pronounced as the state space grows. Our theoretical and empirical analyses relate Value Flattening to an implicit variance penalty in the critic loss and redundant updates from temporally correlated states with similar gradients. Motivated by these findings, we introduce SParse Proximal Policy Optimization (SP$^3$O), which applies the value loss to only a few well-separated states in each response to mitigate both effects. Experiments on Qwen3-Base show that SP$^3$O with only three states supervised per response can mitigate Value Flattening and consistently improve the learned policy across model sizes and evaluation suites. Together, our results identify Value Flattening as an important yet overlooked failure mode of critic learning in standard PPO and show that a simple sparse supervision strategy can mitigate it.