arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03528cs.LGcs.AIcs.AR

LeanGRPO:消除扩散强化学习中的冗余重计算

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

Sijie Wang, Zhiqiang Tan, Xinrui Yang, Shaohuai Shi

首次发表
浏览论文内容

中文总结 AI 辅助

针对扩散强化学习中rollout后冗余重计算的问题,提出LeanGRPO方法,通过两种无重计算训练调度方案实现最高1.83倍端到端加速,且保留原始优化目标。

中文摘要 AI 辅助

扩散强化学习(RL)近期在训练后图像与视频生成模型上取得了显著成功。然而,包括DanceGRPO和FlowGRPO在内的多数扩散RL方法,会在rollout后对选定时间步重新计算并跟踪梯度。在rollout与更新使用相同后端的在线策略训练场景下,这种重计算在数学上是冗余的。直观来看,rollout与策略更新步骤可复用同一前馈骨干网络以避免冗余计算,但此举会在rollout期间产生大量内存开销。为解决该问题,我们提出LeanGRPO,通过重构数据并行布局并引入两种无重计算的轨迹-对数概率扩散RL训练调度方案:(1)LeanGRPO-Retain在rollout期间启用梯度跟踪,直接复用所得计算图与保存的激活值用于更新阶段的反向传播,无需重计算;(2)LeanGRPO-Reweight同样在rollout期间启用梯度,但会使用临时优势立即对每个选定步骤进行反向传播并延迟梯度同步,待轨迹完成后再用真实优势修正临时梯度。这些调度方案适用于不同模型规模与输入尺寸。在基于FLUX.1-dev和Wan的FlowGRPO/DanceGRPO上,LeanGRPO实现了最高1.83倍的端到端加速,同时保留了原始优化目标。

英文摘要

Diffusion reinforcement learning (RL) has recently achieved significant success in post-training image and video generative models. However, most diffusion RL methods, including DanceGRPO and FlowGRPO, recompute selected timesteps with gradient tracking after rollout. Under on-policy training with the same backend for rollout and update, this recomputation is mathematically redundant. Intuitively, the rollout and policy update steps can reuse the same feed-forward backbone to avoid redundant computation, but doing so can incur a large memory overhead during rollout. To address the issue, we present LeanGRPO by restructuring the data-parallel layout and introducing two recompute-free training schedules for trajectory-logprob diffusion RL: (1) LeanGRPO-Retain enables gradient tracking during rollout and directly reuses the resulting computation graphs and saved activations for backward during update, requiring no recomputation; and (2) LeanGRPO-Reweight also enables gradients during rollout, but immediately backpropagates each selected step using a provisional advantage and delays gradient synchronization, then corrects the provisional gradients with the true advantage after the trajectory is completed. These schedules target different model scales and input sizes. Across FlowGRPO/DanceGRPO with FLUX.1-dev and Wan, LeanGRPO achieves up to 1.83x end-to-end speedup while preserving the original optimization objective.

发表机构

  • School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

↑