ABC:基于优势的控制变量用于具有可验证奖励的强化学习
ABC: Advantage-Based Control Variates for Reinforcement Learning with Verifiable Rewards
- MPI-IS(马克斯·普朗克智能系统研究所)
- ELLIS Institute Tübingen(ELLIS 图宾根研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出基于优势的控制变量(ABC)方法,结合直接优势估计(DAE)形成演员-评论家算法,在离线到在线RLVR设置中,以更少在线步骤达到与GRPO相当的性能。
AI中文摘要:
近年来,具有可验证奖励的强化学习(RLVR)的进展凸显了如组相对策略优化(GRPO)等简单无评论家策略梯度方法的有效性。相比之下,演员-评论家方法依赖于学习到的价值函数,其近似误差可通过常用的优势估计器(如时序差分误差)引入偏差。受此观察启发,我们通过优势-价值公式重新审视轨迹级控制变量,称之为基于优势的控制变量(ABC)。该公式揭示了协方差结构与直接优势估计(DAE)中使用的回报分解密切相关。最后,我们将ABC与DAE结合成一个单一的演员-评论家算法,并在离线到在线RLVR设置中进行评估,其中评论家首先在先前收集的轨迹上训练,并在在线学习期间适应。在数学推理任务上,ABC在使用显著更少的在线优化步骤的情况下实现了与GRPO相当的性能。
英文摘要:
Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient methods such as Group Relative Policy Optimization (GRPO). In contrast, actor-critic methods rely on learned value functions whose approximation error can introduce bias through commonly used advantage estimators such as temporal-difference error. Motivated by this observation, we revisit trajectory-level control variates through an advantage-value formulation, which we call Advantage-Based Control Variates (ABC). This formulation reveals that the covariance structure is closely related to the return decomposition used in Direct Advantage Estimation (DAE). Finally, we combine ABC with DAE into a single actor-critic algorithm and evaluate it in an offline-to-online RLVR setting, where the critic is first trained on previously collected trajectories and adapted during online learning. On mathematical reasoning tasks, ABC achieves performance competitive with GRPO using substantially fewer online optimization steps.