arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36932cs.AI

从差距中学习:基于自适应滚动采样的差异感知优势剪枝用于GRPO

Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO

Jiahua Yang, Zhiwei Yang, Xianpeng Zhang, Dongyu Chen, Xing Chen, Tianhuang Su, Haonan Lu, Quanlong Guan, Kai Tang, Chuangchuang Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对GRPO训练中计算开销大和低信息轨迹影响学习效率的问题,提出FastRL框架,通过优势感知剪枝和自适应滚动采样,在多个基准上实现2.07倍加速和1.64%准确率提升。

中文摘要 AI 辅助

近年来,群体相对策略优化(GRPO)及其变体已被开发用于策略优化,并展现出显著的性能提升。然而,这些方法通常因每个问题多次滚动采样以及跨滚动重复的逐标记概率评估而产生大量计算开销。此外,低信息或高度同质的轨迹会降低下游学习信号的效率,阻碍模型优化并限制最终性能。为解决这些问题,我们提出了FastRL,一种新颖的即插即用强化学习框架,同时提高训练效率和策略学习效果。具体而言,1)我们引入了一种优势感知剪枝策略,选择性地保留高优势轨迹,同时最大化轨迹间的梯度多样性。2)然后,我们设计了一种自适应滚动采样机制,根据历史剪枝分布动态调整不同训练阶段的采样规模,在探索充分性和计算效率之间取得平衡。实验表明,FastRL可以无缝集成到GRPO、DAPO和GSPO变体中,在Geometry3K和GeoQA8K-R1V上实现了平均2.07倍的训练加速,并在视觉推理基准上平均准确率提升约1.64%。源代码将在该https URL上提供。

英文摘要

Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning framework that simultaneously improves training efficiency and the effectiveness of policy learning. Specifically, 1) We introduce an advantage-aware pruning strategy to selectively preserve high-advantage trajectories while maximizing inter-trajectory gradient diversity. 2) Then, we design an adaptive rollout sampling mechanism to dynamically adjust the sampling scale across different training stages based on historical pruning distributions, balancing exploration adequacy and computational efficiency. Experiments demonstrate that FastRL can be seamlessly integrated into GRPO, DAPO, and GSPO variants, achieving an average 2.07$\times$ training speedup on Geometry3K and GeoQA8K-R1V, along with an approximately 1.64\% improvement in average accuracy on visual reasoning benchmarks. Source codes will be available at https://github.com/Nicozwy/FastRL.

发表机构

  • Guangdong Institute of Smart Education, Jinan University(暨南大学广东智慧教育研究院)
  • OPPO AI Center(OPPO AI中心)
  • Ragentile Intelligence Inc(Ragentile Intelligence公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑