arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习剩余内容,而非已掌握内容:面向多奖励策略优化的饱和感知优势重加权方法

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

Yixuan Wang, Yifei Chen, Haichao Zhang, Haozheng Luo, Xander Wu, Jie Ni, Yun Fu, Nuno Vasconcelos, Yijiang Li

arXiv 2608.16072首次发表:更新:

发表机构

University of Florida; UC San Diego; Northeastern University; Northwestern University; Stanford University; Universität Innsbruck(佛罗里达大学; 加州大学圣迭戈分校; 东北大学; 西北大学; 斯坦福大学; 因斯布鲁克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多奖励策略优化中现有方法的缺陷,提出SA-MRPO方法,动态分配优化资源,在数学推理、自适应推理、编码任务中均实现性能提升。

AI 中文摘要

采用组相对优势的强化学习(RL)已成为训练后语言模型推理器的事实上的标准。然而,在优化多个奖励目标时,现有方法通常在组内标准化之前,用固定加权和对奖励向量进行标量化。我们证明这种设计会导致两个根本问题:具有不同奖励分布的轨迹会获得相同的优势,且所有目标都以固定相对权重进行优化,无论其当前饱和程度如何。结果,训练过程会持续将梯度预算分配给已解决的目标,而非聚焦于剩余提升空间更大的目标。我们提出多奖励策略优化的饱和感知优势重加权方法(SA-MRPO),该方法对每个奖励目标独立进行标准化,并根据目标饱和程度的批次级估计自适应地降低其贡献权重。这会动态地将优化资源重新分配给欠优化的目标,同时在经验上维持已充分满足的目标的性能。我们进一步表明,饱和感知重加权可以反转更新的符号,而非仅缩放其幅度。在包含两目标和三目标奖励组合的数学推理任务中,SA-MRPO在15次基准比较中有12次比GDPO提升了更难的正确性目标,在AIME24上的提升最高达5%;在自适应推理任务中,其在所有5个基准上均提升了准确率,平均提升3.8%,在AMC23上的提升最高达9.2%;在编码基准任务中,其通过率最高提升2.3%,同时在所有设置中均将较易目标维持在已充分满足的水平附近。

英文摘要

Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that this design leads to two fundamental problems: rollouts with distinct reward profiles can receive identical advantages, and all objectives are optimized with fixed relative weights regardless of their current level of saturation. Consequently, an objective whose rewards are already near their upper bound can retain substantial influence when its rewards still vary within rollout groups. We introduce Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation. This changes the relative contribution of each objective according to its observed reward headroom. We further derive an exact condition under which saturation aware reweighting reverses the sign of a rollout's aggregate advantage. Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to $5\%$ on AIME24. On adaptive reasoning it improves accuracy on all five benchmarks, by $3.8\%$ on average and up to $9.2 \%$ on AMC23, and on coding benchmarks it improves pass rate by up to $2.3\%$, while in all settings maintaining the easier objectives near their already satisfied levels. Additional experiments characterize the accompanying reward tradeoffs and sensitivity to corrupted rewards.

Comments18 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑