贝尔曼策略优化
Bellman Policy Optimization
浏览论文内容
中文总结 AI 辅助
本文提出贝尔曼策略优化(BPO),一种基于策略镜像下降的无评论家强化学习方法,通过贝尔曼方程将目标转化为轨迹级形式,避免中间状态值估计,并在数学推理基准上验证了其有效性。
中文摘要 AI 辅助
具有可验证奖励的强化学习(RLVR)提升了大型语言模型(LLMs)的推理能力。我们引入了贝尔曼策略优化(BPO),这是一种源自策略镜像下降(PMD)的无评论家方法。对于具有终止奖励的自回归生成,BPO利用贝尔曼方程将PMD重新表述为轨迹级目标。该重新表述避免了对中间状态的状态值估计。我们证明了它与原始PMD目标具有相同的唯一最优解。我们通过近似该目标推导出实用的BPO损失。其失配校正权重是互补令牌概率的平滑比率。在数学推理基准上的实验证明了BPO的有效性。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs). We introduce Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD). For autoregressive generation with terminal rewards, BPO uses the Bellman equations to reformulate PMD as a trajectory-level objective. The reformulation avoids estimating state values at intermediate states. We prove that it has the same unique optimal solution as the original PMD objective. We derive the practical BPO loss by approximating this objective. Its mismatch-correction weight is a smoothed ratio of complementary token probabilities. Experiments on mathematical reasoning benchmarks demonstrate the effectiveness of BPO.
发表机构
- Apodex US, Inc.(Apodex US公司)
- Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。