arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00444cs.LGcs.CL

分组自适应裁剪策略优化

Group Adaptive Clipping Policy Optimization

Sheng Jia, Xiao Wang, Shiva Prasad Kasiviswanathan, Rein Houthooft

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对GRPO的固定裁剪局限,提出GAPO方法,通过自适应调整裁剪边界提升了Qwen、Llama模型在低通过率数学推理与编码任务上的Pass@1和Pass@k指标。

中文摘要 AI 辅助

针对可验证奖励强化学习(RLVR)的分组相对策略优化(GRPO)通常在所有rollout中使用固定的重要性采样(IS)比率裁剪边界。我们发现一个关键局限:更难问题上的罕见正确rollout与更简单问题上的大量正确rollout的裁剪率相近,尽管它们贡献的学习信号差异极大。分组成功率低的rollout具有更大的IS比率,携带更强的探索和解决新问题的梯度信号,但会被固定裁剪不成比例地抑制。为解决该问题,我们提出分组自适应裁剪策略优化(GAPO),这是对GRPO方法的插件式修改,可根据rollout优势自适应调整裁剪边界。GAPO基于反向KL信任域视角,该视角表明具有更大学习信号的rollout应获得成比例更大的更新空间。GAPO无需奖励塑形,保留标准PPO/GSPO代理,仅调整裁剪阈值。在Qwen和Llama模型上,当基础模型的通过率相对较低时,GAPO在数学推理和编码基准上,始终比固定裁剪和优势塑形基线提高Pass@1和Pass@k。

英文摘要

Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed importance-sampling (IS) ratio clipping boundary across all rollouts. We identify a key limitation: rare correct rollouts on harder problems and abundant correct rollouts on easier problems are clipped at comparable rates, despite contributing very different learning signals. Rollouts with low group success exhibit larger IS ratios and carry stronger gradient signal for exploration and solving new problems, yet are disproportionately suppressed by fixed clipping. To address this, we propose Group Adaptive Clipping Policy Optimization (GAPO), a plug-in modification to GRPO methods that adapts the clipping boundary to the rollout advantage. GAPO is motivated by a reverse-KL trust-region perspective, which suggests that rollouts with larger learning signal should receive proportionally greater update headroom. GAPO requires no reward shaping and preserves the standard PPO/GSPO surrogate while adapting only the clipping threshold. Across Qwen and Llama models, GAPO consistently improves both Pass@1 and Pass@k over fixed clipping and advantage-shaping baselines on math reasoning and coding benchmarks where the pass rates by the base model are relatively low.

发表机构

  • University of Toronto(多伦多大学)
  • Amazon(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑