arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ORPG:通过目标级策略梯度调和多个奖励目标

ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients

Shicheng Fang, Yiwen Zhao, Wenbo Tian, Jiahao Lu, Yining Zheng, Yuxin Wang, Xipeng Qiu

arXiv 2609.34985首次发表:更新:

发表机构

Fudan University; Shanghai Innovation Institute(复旦大学; 上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ORPG通过目标级梯度调和实现多奖励策略优化,在有用性-安全性和数学推理任务中显著提升性能,支持兼容与冲突梯度的统一处理。

AI 中文摘要

多奖励策略优化需要一种联合更新,该更新既要反映学习信号,也要反映各目标之间的预期关系。我们引入了目标级调和策略梯度(ORPG),该方法为每个奖励构建一个独立的裁剪策略目标,并将由此产生的梯度调和为一次策略更新。对于兼容的梯度,一种基于余弦的插值通过部分归一化的参考来协调它们的贡献,同时保持其和的范数。我们将此更新刻画为球形方向折中的唯一解。对于冲突的梯度,投影遵循任务的优先级。我们在有用性-安全性对齐以及数学推理的正确性-成本优化中评估了相同的兼容规则。ORPG在每条轴上相较于最强的外部基线显著提高了平均有用性和无害性得分。在数学任务中,它在所比较的方法中实现了最高的平均全预算准确率和三预算超体积,并且其响应比初始策略更准确、更简短。组件比较和训练动态表明,兼容协调贡献更大,而冲突处理则提供了互补的收益。这些结果支持对地位平等的目标以及具有明确优先级的目标进行梯度调和。

英文摘要

Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients into one policy update. For compatible gradients, a cosine-dependent interpolation coordinates their contributions through a partially normalized reference while preserving the norm of their sum. We characterize this update as the unique solution of a spherical directional compromise. For conflicting gradients, projection follows the task's priorities. We evaluate the same compatible rule in helpfulness--safety alignment and correctness--cost optimization for mathematical reasoning. ORPG substantially improves average Useful and Harmless scores over the strongest external baseline on each axis. In mathematics, it achieves the highest average full-budget accuracy and three-budget hypervolume among the compared methods, with more accurate and shorter responses than the initial policy. Component comparisons and training dynamics show the larger contribution of compatible coordination and a complementary benefit from conflict handling. These results support gradient reconciliation for objectives with equal standing and for objectives with an explicit priority.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑