发表机构
Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen); Beijing Zhongguancun Academy; XinzhuAI(哈尔滨工业大学(深圳)计算与智能学院; 北京中关村学院; 芯智科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对GRPO等RL方法的策略熵崩溃问题,提出GRPODropout方法,通过选择性移除少量高概率正优势rollout提升性能,仅产生可忽略的计算开销,实验验证其在更少样本下仍能取得更优效果。
AI 中文摘要
GRPO等强化学习(RL)方法大幅提升了大语言模型的推理能力,但常出现策略熵崩溃问题:采样多样性的丧失会削弱探索能力,限制模型的进一步提升。现有方法通过算法层面的干预(如奖励修改、熵/KL正则化)或token层面的重加权来解决该问题。本文从互补视角展开研究:熵崩溃也可通过调整哪些生成的rollout参与策略更新来缓解。在相同采样预算下,并非所有rollout都对更新有正向贡献,有选择性地排除部分rollout可提升学习效果。为此,本文提出GRPODropout:在标准更新前,采用简单策略选择性移除少量高概率正优势rollout,并对保留的优势值进行重新中心化。为支撑该设计,本文开展了rollout层面的理论分析,为方法设计和阈值选择提供指导。该方法仅调整rollout的使用方式,计算开销可忽略不计。实验结果表明,与原始GRPO相比,该方法在更新时使用更少的rollout样本,却能获得更高的准确率和更高的策略熵,印证了“少即是多”的理念。本研究为RL rollout的使用提供了新的见解:移除部分rollout可提升性能。代码可在指定URL获取。
英文摘要
Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such as reward modification and entropy/KL regularization, or through token-level reweighting. We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates. Under the same sampling budget, not all rollouts contribute positively to an update, and selectively excluding some can improve learning. To address this, we propose GRPODropout: before the standard update, we use a simple strategy that selectively removes a small number of high-probability positive-advantage rollouts and recenters the retained advantages. To motivate this design, we develop a rollout-level theoretical analysis that guides method design and threshold selection. The method changes only rollout usage, and adds negligible computational overhead. Experiments show higher accuracy than original GRPO and higher actor entropy while using fewer rollout samples for updates, illustrating "less is more." This work provides insight into RL rollout usage: removing some rollouts can improve performance. Code is available at https://github.com/hexuandeng/GRPODropout/.