发表机构
The Ohio State University; Arizona State University(俄亥俄州立大学; 亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出异步GRPO的收敛分析,并设计GMC-GRPO方法通过组质量上限最小化偏差,改善延迟项阈值依赖性,实验验证其鲁棒性。
AI 中文摘要
异步强化学习(RL)提高了大语言模型后训练的效率,但引入了由较早策略生成的陈旧轨迹。关于这种陈旧性如何影响收敛性以及如何减轻其影响的理论理解仍然有限。我们为GRPO风格的算法推导了一个收敛界,该界明确刻画了梯度估计器的二阶矩与偏差之间的权衡。对于轨迹级重要性加权估计器,我们的分析表明,一旦二阶矩被均匀控制,延迟通过裁剪或重新缩放引入的偏差进入界中。受此洞察指导,我们提出了一种新颖的组质量上限GRPO(GMC-GRPO)方法,该方法在一类共享共同二阶矩保证的加权估计器中最小化基于比率的偏差界。我们为异步GMC-GRPO建立了收敛保证,并表明与TIC-GRPO相比,当ε→0时,它将四阶延迟项的阈值依赖性从O(ε^{-4})改善到O(ε^{-2}),其中1+ε是比率阈值。在局部策略重叠下,调整步长后,延迟相关项随G^{-2/5}减少,其中G是组大小。对于固定的行为策略和当前策略,组重新缩放引入的偏差也随G→∞而消失,而轨迹级裁剪的偏差可能持续存在。在Qwen3模型和推理基准上的实验表明,对陈旧轨迹的鲁棒性有所提高,GMC-GRPO在大延迟下在稳定基线中取得了最佳性能。
英文摘要
Asynchronous reinforcement learning (RL) improves the efficiency of large language model post-training but introduces stale rollouts generated by earlier policies. Theoretical understanding of how this staleness affects convergence and how to mitigate its impact remains limited. We derive a convergence bound for GRPO-style algorithms that explicitly characterizes the tradeoff between the gradient estimator's second moment and bias. For trajectory-level importance-weighted estimators, our analysis shows that once the second moment is uniformly controlled, delay enters the bound through the bias introduced by clipping or rescaling. Guided by this insight, we propose a novel group mass capping GRPO (GMC-GRPO) method, which minimizes a ratio-based bias bound within a class of weighted estimators sharing a common second-moment guarantee. We establish convergence guarantees for asynchronous GMC-GRPO and show that, compared with TIC-GRPO, it improves the threshold dependence of the fourth-order delay term from $O(ε^{-4})$ to $O(ε^{-2})$ as $ε\to0$, where $1+ε$ is the ratio threshold. Under local policy overlap, the delay-dependent term decreases as $G^{-2/5}$ after tuning the step size, where $G$ is the group size. For fixed behavior and current policies, the bias introduced by group rescaling also vanishes as $G\to\infty$, whereas the bias from trajectory-wise clipping can persist. Experiments across Qwen3 models and reasoning benchmarks demonstrate improved robustness to stale rollouts, with GMC-GRPO achieving the best performance among stable baselines under large rollout delays.
Comments40 pages, 6 figures