发表机构
Fudan University; Shanghai Innovation Institute; Tencent YoutuLab; Griffith University; Shanghai Academy of Artificial Intelligence for Science(复旦大学; 上海创新研究院; 腾讯优图实验室; 格里菲斯大学; 上海科学人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出HL-Gauss PPO,将PPO的标量评判器改为分类式预测器,在多任务及Qwen2.5、Qwen3主干上均优于PPO、DAPO基线,可改进评判器信号与优势估计。
AI 中文摘要
针对大语言模型的近端策略优化(PPO)算法通常通过对标量价值目标进行均方误差(MSE)回归来训练其评判器。尽管标量MSE在统计上对估计条件期望回报是有效的,但在带可验证奖励的强化学习(RLVR)中,稀疏二元奖励使得评判器的优化与校准尤为关键:微小的价值误差会直接扭曲PPO所用的标量优势。我们研究基于分类的训练目标是否可改进该评判器信号。HL-Gauss PPO将标量MSE头替换为对离散化价值支持的分类预测器,通过针对平滑HL-Gauss目标的交叉熵进行训练;其输出被解码为标准广义优势估计(GAE)和PPO的标量期望,因此策略更新保持不变,并非分布式。在数学推理、工具增强数学及Search-R1任务中,基于Qwen2.5和Qwen3两种主干,HL-Gauss PPO始终优于强基线PPO和DAPO。对独热、双热及伯努利双箱评判器的对照实验表明,更大的输出头或仅二元分类均无法解释性能提升。在通用推理前缀集上,HL-Gauss降低了布雷分数和校准误差,产生更对称、方差更低的优势。这些结果表明,分类式价值学习是RLVR中PPO评判器的有效优化替代方案。
英文摘要
Proximal Policy Optimization (PPO) for large language models typically trains its critic by mean-squared-error (MSE) regression on scalar value targets. Although scalar MSE is statistically valid for estimating the conditional expected return, sparse binary rewards in reinforcement learning with verifiable rewards (RLVR) make critic optimization and calibration especially consequential: small value errors directly distort the scalar advantages used by PPO. We study whether a classification-based training objective can improve this critic signal. HL-Gauss PPO replaces the scalar MSE head with a categorical predictor over a discretized value support, trained by cross-entropy against smoothed HL-Gauss targets. Its output is decoded to a scalar expectation for standard GAE and PPO; the actor update is therefore unchanged and is not distributional. Across mathematical reasoning, tool-augmented math, and Search-R1, and on both Qwen2.5 and Qwen3 backbones, HL-Gauss PPO consistently improves over strong PPO and DAPO baselines. Controls with one-hot, two-hot, and Bernoulli two-bin critics show that neither a larger output head nor binary classification alone explains the gains. On a common collection of reasoning prefixes, HL-Gauss improves Brier score and calibration error and yields more symmetric, lower-variance advantages. These results position categorical value learning as an effective optimization surrogate for PPO critics in RLVR.
CommentsAccepted at COLM 2026. 26 pages, 9 figures. Code: https://github.com/ZhijianZhou/HL-guass-ppo