arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SMOPD:通过“专门化-合并”在线策略蒸馏实现多奖励强化学习

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation

Wen Wang, Jiahua Bao, Tu Yongsiqi, Yihao Liu, Haotian Zhou, Haoxuan Ma, Mengyu Zhou, Wenkui Fan, Junwei He, Xiaoxi Jiang, Guanjun Jiang

arXiv 2608.03092首次发表:更新:

AI 中文总结

针对多奖励强化学习中GDPO难以平衡不同粒度奖励信号的问题,提出SMOPD两阶段训练方法,在1.5B、3B、7B骨干模型上的互补与冲突奖励设置下性能优于GDPO。

AI 中文摘要

我们旨在提升多奖励强化学习训练过程中的模型性能。现有分组奖励解耦归一化策略优化(GDPO)已通过在聚合前单独归一化每个奖励维度,缓解了直接标量化过程中奖励信号相互掩盖的问题。然而,我们的实验表明,GDPO仍难以平衡粒度不同的奖励信号。具体而言,在某些特定训练任务中,模型可能同时收到两种奖励:一种是赋予0.1至1.0范围内细粒度分数的密集奖励,另一种是仅提供0或1二元反馈的稀疏奖励。在这种情况下,我们发现稀疏奖励可能提供不足的优化信号,导致其对应的能力无法得到有效强化。因此,如何在不牺牲已从细粒度奖励中学习到的能力的前提下,增强稀疏奖励的优化信号?为克服这一局限,我们提出“专门化-合并”在线策略蒸馏(SMOPD),这是一种用于多奖励优化的两阶段训练方法。第一阶段:专门化(Stage1-Specialize):SMOPD首先采用奖励优先级配置训练多个奖励专门化的教师模型,使每个奖励在其信号能有效驱动优化的条件下进行学习。第二阶段:合并(Stage2-Merge):SMOPD随后利用在线策略蒸馏将这些教师模型的奖励专门化能力合并为单一学生策略,同时维持任务级优化的平衡。为验证我们的方法,我们在两种多奖励设置下开展实验:互补奖励(工具调用准确率与格式)和冲突奖励(有用性与无害性奖励)。基于上述设置,SMOPD在15亿(1.5B)、30亿(3B)和70亿(7B)参数规模的骨干模型上均优于GDPO。

英文摘要

We aim to improve model performance in multi-reward reinforcement learning training process. Existing Group reward-Decoupled Normalization Policy Optimization (GDPO) has mitigated the issue of reward signals masking one another during direct scalarization by normalizing each reward dimension separately before aggregation. However, our experiments show that GDPO still struggles to balance reward signals with different granularities. Specifically, in some particular training tasks, the model may receive a dense reward that assigns fine-grained scores ranging from 0.1 to 1.0, together with a sparse reward that provides only binary feedback of either 0 or 1. In such cases, we find that the sparse reward may provide an insufficient optimization signal, preventing its corresponding capability from being effectively reinforced. Therefore, how can we strengthen the optimization signal from the sparse reward without sacrificing the capability already learned from the fine-grained reward? To overcome this limitation, we propose Specialize-and-Merge Online Policy Distillation (SMOPD), a two-stage training method for multi-reward optimization. Stage1-Specialize: SMOPD first employs reward-priority configurations to train multiple reward-specialized teachers, allowing each reward to be learned under conditions where its signal can effectively drive optimization. Stage2-Merge: SMOPD then utilizes online policy distillation to combine the reward-specialized capabilities of these teachers into a single student policy, while maintaining balanced task-level optimization. To validate our method, we conduct experiments on two multi-reward settings: complementary rewards(tool-calling accuracy and format) and conflicting rewards (helpful and harmless rewards). Based on above settings, SMOPD outperforms GDPO across 1.5B, 3B and 7B backbones.

Comments21 pages, 5 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑