发表机构
University of Chinese Academy of Sciences; University of Minnesota Twin Cities; University of Wisconsin–Madison; Stony Brook University; Dalian University of Technology; Beijing Foreign Studies University; Southeast University; Kuaishou Technology(中国科学院大学; 明尼苏达大学双城分校; 威斯康星大学麦迪逊分校; 纽约州立大学石溪分校; 大连理工大学; 北京外国语大学; 东南大学; 快手科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多奖励强化学习中奖励稀疏导致的学习不均衡问题,提出密度感知奖励聚合(DARA),通过逆平方根密度校正加权奖励信号,加速工具调用与数学推理任务的学习,减少训练步数。
AI 中文摘要
多奖励强化学习训练大型语言模型同时满足多个行为目标。GDPO中使用的按奖励归一化保留了rollout组内特定于奖励的相对信息,但不同目标仍可能表现出不均衡的学习进度。我们通过优势能量(即一个奖励在一个批次内平方优势之和)来研究这种行为。在理想化的GDPO归一化下,我们证明该能量与活跃组密度成正比,即奖励在rollout组中提供非零相对优势的组所占比例。这揭示了残余的批次级信号不平衡,并为校准奖励贡献提供了基础。基于这一关系,我们提出了密度感知奖励聚合(DARA)。我们推导出逆平方根密度校正,该校正赋予较不频繁活跃的奖励信号更大的权重。DARA从每个rollout批次计算其权重,适应训练过程中奖励活跃度的变化,而不修改底层策略优化目标。在工具调用和数学推理上的实验表明,DARA比GDPO更快地学习目标行为,在工具调用上达到高格式合规性所需的训练步数最多减少26%,在数学推理上达到接近饱和的长度合规性所需的步数最多减少65%,同时在最终性能上保持竞争力。我们的代码可在以下网址获取:此https URL。
英文摘要
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at https://github.com/zhaihaotian/DARA.
Comments24 pages, 8 figures