双通道鲁棒群相对策略优化:基于优势与序列权重估计
Dual-Channel Robust Group-Relative Policy Optimization via Advantage and Sequence-Weight Estimation
- Beihang University(北京航空航天大学)
- Peking University(北京大学)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对群相对策略优化对异常值敏感的问题,提出双通道鲁棒优化器RoVR-GSPO,分别处理奖励与比率通道,实验证明其在多个任务上优于GSPO且鲁棒性更强。
AI中文摘要:
群相对策略优化依赖于奖励衍生的优势与序列级似然权重,这两者都可能对局部异常值敏感。极端奖励会在群归一化后压缩干净响应之间的对比度,而令牌级对数比率扰动则可能改变序列权重和裁剪决策。我们提出RoVR-GSPO,一种双通道鲁棒优化器,分别解决这些失败模式。其奖励通道结合了鲁棒参考估计与有界残差信用,而比率通道使用可微的SoftRoVR聚合来构建鲁棒序列权重。我们提供了两个通道的稳定性与效率分析。在数学推理、长上下文摘要和工具调用标注上的实验显示,与GSPO相比有持续改进,而受控扰动研究则证明了对奖励污染和令牌比率异常的更强鲁棒性。
英文摘要:
Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while token-level log-ratio perturbations can alter sequence weights and clipping decisions. We introduce RoVR-GSPO, a dual-channel robust optimizer that addresses these failure modes separately. Its reward channel combines robust reference estimation with bounded residual credit, while its ratio channel uses differentiable SoftRoVR aggregation to construct robust sequence weights. We provide stability and efficiency analyses for both channels. Experiments on mathematical reasoning, long-context summarization, and tool-call annotation show consistent improvements over GSPO, while controlled perturbation studies demonstrate stronger robustness to reward contamination and token-ratio anomalies.