窥见箔片:利用自特权评论家重塑价值估计以用于 RLVR
Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR
浏览论文内容
中文总结 AI 辅助
针对 LLM 多步推理中信用分配困难,本文提出自特权演员-评论家框架 πPPO,利用已验证轨迹作对比证据提升价值估计质量,在数学推理基准上超越现有基线。
中文摘要 AI 辅助
在具有稀疏终端奖励的多步推理任务中训练大型语言模型(LLM)时,为中间步骤分配信用仍然是一个核心挑战,而诸如 PPO 之类的演员-评论家方法通过学习价值函数来构建词元级优势来解决这一问题。然而,其有效性依赖于可靠的价值估计,这是一项困难的任务,要求评论家既要评估朝向正确解决方案的进展,又要预测不断演化的策略的未来行为;任一方面的错误都可能损害信用分配并使在线训练不稳定。在本文中,我们重新审视了标准的仅状态形式的价值估计,并提出了 πPPO,一个自特权的演员-评论家框架。通过重用经过验证的相同提示的轨迹作为对比证据,πPPO 帮助评论家对照成功和失败的尝试来评估中间推理,同时保留标准的策略优化和部署接口。实验表明,πPPO 在价值估计质量上持续大幅提升,并在具有挑战性的数学推理基准上优于代表性的演员-评论家和无评论家的 RLVR 基线,即使与规模显著更小的非对称评论家配对时也保持有效。
英文摘要
Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges on reliable value estimation, a difficult task requiring the critic to both assess progress toward a correct solution and anticipate an evolving policy's future behavior; errors in either can compromise credit assignment and destabilize online training. In this paper, we revisit the standard state-only formulation of value estimation and propose $π$PPO, a self-privileged actor-critic framework. By reusing verified same-prompt rollouts as contrastive evidence, $π$PPO helps the critic assess intermediate reasoning against successful and failed attempts, while preserving standard policy optimization and the deployment interface. Experiments show that $π$PPO consistently improves value-estimation quality by a substantial margin and outperforms representative actor-critic and critic-free RLVR baselines on challenging mathematical reasoning benchmarks, while remaining effective even when paired with substantially smaller asymmetric critics.
发表机构
- Peking University(北京大学)
- Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。