arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.01083cs.LGcs.AI

异步RLHF的陈旧度-学习率缩放定律

Scaling Laws for Collapse in Asynchronous GRPO

  • The University of Hong Kong(香港大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Gradient
  • University of Southern California(南加州大学)
  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Jingwei Song, Haofeng Xu, Jie Xiao, Chengke Bao, Jingwei Shi, Pengbin Feng, Yuhang Han, Weixun Wang, Eric Yang, Tianyu Shi

中文总结 AI 辅助

研究异步GRPO中陈旧rollout的影响,推导出陈旧度与学习率之间的缩放定律,揭示稳定性受累积漂移和陈旧度约束的双重条件。

中文摘要 AI 辅助

高吞吐量RLHF系统通常将rollout生成与策略优化解耦,导致学习器更新时使用陈旧rollout。本文研究异步GRPO中这种陈旧性的影响。我们在GRPO代理目标中显式引入行为策略,并区分学习器使用的代理梯度映射与分布依赖种群目标的真实全导数。在局部有界性、分布平滑性和行为策略平滑性假设下,我们证明陈旧rollout引入每步代理梯度偏差,阶数为O(S * eta),其中S表示最大rollout滞后,eta表示学习率。我们进一步推导出条件性崩溃时间缩放定律:当周期内漂移低于批次级裁剪半径时,崩溃主要由累积学习器漂移T * eta控制;当陈旧rollout约束激活时,稳定性显式依赖于S * eta。这产生双约束稳定性条件eta << min{R_batch / (S * G_upd), R_crit / (T * G_upd)},解释了在有限视野范围内最大稳定学习率为何可能对陈旧性呈现弱依赖。

英文摘要

Asynchronous reinforcement learning improves the throughput of large language model post-training by decoupling rollout generation from policy optimization, but introduces a mismatch between the behavior and learner policies. How the resulting policy staleness couples with the learning rate to govern training stability and collapse time remains poorly understood. We investigate this coupling in vanilla GRPO through controlled sweeps of the synchronization interval $S$ and constant learning rate $η$ on Llama-3.2-1B/3B, complemented by experiments on Qwen3-8B. We identify two empirical scaling laws: (i) Stability-boundary scaling: the largest stable learning rate scales approximately as $S^{-1}$, yielding a stability boundary characterized by an approximately constant product $Sη$. (ii) Collapse-time scaling: among collapsing runs, estimated collapse times scale approximately as $η^{-1}$, corresponding to a model- and setup-dependent cumulative learning-rate budget that aligns across synchronization intervals in the Llama sweeps. We interpret these laws through a local analysis of the behavior-dependent GRPO surrogate and a complementary mean-field model. Under local regularity conditions, the analysis yields an $O(Sη)$ upper bound on the staleness-induced update bias that resets at synchronization. The mean-field model shows how sufficiently strong positive feedback can sustain directional drift when update directions persist across synchronization cycles. When drift speed saturates under optimizer normalization, this mechanism predicts exit from a local surrogate-validity region after an approximately fixed cumulative learning rate. Together, these findings motivate a practical calibration rule: estimate the stability threshold and collapse budget from a coarse sweep, then jointly select $S$ and $η$ for the intended training horizon.

补充信息

↑