发表机构
National University of Singapore; Tencent Hunyuan(新加坡国立大学; 腾讯混元)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出BPCO流程稳定高效训练评价模型,在数学推理任务上,其优于基线,且采样1个响应时可匹配或超过基于组的GRPO基线。
AI 中文摘要
针对大型语言模型的基于组的强化学习方法(如GRPO)通过为每个提示采样多个响应来避免训练评价模型。可靠的评价模型可从单个响应中估计token级优势,但标准的基于评价模型的训练流程往往不稳定。我们研究该不稳定性并提出最佳实践评价模型优化(BPCO),该流程结合DPPO、限定在奖励范围内的价值预测、蒙特卡洛价值目标、未归一化的策略优势以及长度自适应的广义优势估计。由于评价模型仅在训练期间使用,BPCO还可使其依赖于奖励定义信息(如参考答案或评分标准),这些信息对策略隐藏。受控实验分离了每个设计选择的效果。在数学推理任务上,模型规模从15亿参数到30B-A3B混合专家模型,BPCO始终提升了强大的基于评价模型的基线,且每个提示采样1个响应时可匹配或超过基于组的基线。相同流程也提升了基于评分标准的奖励学习。这些结果表明,精心设计的评价模型是组相对优势估计的可靠替代方案。代码可在该https URL获取。
英文摘要
Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token-level advantages from one response, but standard critic-based training recipes are often unstable. We study this instability and develop **Best Practice Critic Optimization (BPCO)**, a recipe that combines DPPO, value predictions bounded to the reward range, Monte Carlo value targets, unnormalized policy advantages, and length-adaptive generalized advantage estimation. Because the critic is used only during training, BPCO can also condition it on reward-defining information, such as a reference answer or grading rubric, that is hidden from the policy. Controlled experiments isolate the effect of each design choice. Across mathematical reasoning tasks with models ranging from 1.5B parameters to 30B-A3B mixtures of experts, BPCO improves a strong critic-based baseline consistently, and matches or exceeds a group-based baseline while sampling one response per prompt. The same recipe also improves learning with rubric-based rewards. These results show that a carefully designed critic provides a reliable alternative to group-relative advantage estimation. Code is available at https://github.com/QPHutu/golden_critic.