arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35433cs.LGcs.AI

ReSPO:用于离线策略学习中梯度饥饿问题的重塑序列策略优化

ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning

  • University of California, Los Angeles(加州大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

Yihang Chen, Yuanhao Ban, Cho-Jui Hsieh

AI总结:

针对离线策略学习中裁剪导致的梯度饥饿问题,提出ReSPO方法,用平滑双分支核替代裁剪,在Qwen3模型上加速优化并提升最终性能。

AI中文摘要:

基于可验证奖励的强化学习(RLVR)经常在多次策略更新中重复使用rollout数据,这加剧了当前策略与数据生成策略之间的不匹配。我们在裁剪策略优化中发现了一个依赖于符号的梯度饥饿问题:裁剪在低重要性权重尾部抑制了生成不足的正响应,同时却允许严重过度生成的负响应主导高权重尾部。为解决此问题,我们提出了ReSPO(重塑序列策略优化),该方法用从α-散度变分目标和指数方差控制倾斜导出的平滑双分支序列级核替代裁剪。正分支为生成不足的正响应保留非零梯度权重,而负分支则抑制严重过度生成的负响应。我们证明,即使在累积的策略漂移将长正推理轨迹降级到低重要性权重尾部时,ReSPO也能在早期训练中有效地从这些轨迹中学习。在密集和MoE Qwen3模型上,ReSPO加速了早期优化,提高了最终训练分数,并在rollout重用下取得了更高的保留基准性能,验证了我们在离线策略学习中对重要性权重尾部控制的方法。

英文摘要:

Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate the high-weight tail. To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an $α$-divergence variational objective and an exponential variance-control tilt. The positive branch preserves a nonzero gradient weight for under-generated positive responses, while the negative branch suppresses heavily over-generated negative responses. We demonstrate that ReSPO effectively learns from long positive reasoning trajectories during early training, even when accumulated policy drift relegates them to the low-importance-weight tail. On dense and MoE Qwen3 models, ReSPO accelerates early optimization, improves final training scores, and achieves higher held-out benchmark performance under a rollout reuse, validating our approach on importance-weight tail control in off-policy learning.

↑