arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36816cs.LG

迈向更好的训练信号:优势裁剪策略优化

Towards Better Training Signal: Advantage Clipped Policy Optimization

Ruichuan Huang, Jinghan Liu, Congliang Chen

首次发表
浏览论文内容

中文总结 AI 辅助

提出ACPO算法,通过裁剪重要性采样比率与优势的乘积来稳定梯度估计,在数学推理基准上较PPO和GRPO提升4-6个百分点,并证明其与策略镜像下降中梯度裁剪的关联及收敛性。

中文摘要 AI 辅助

强化学习(RL)已成为提升大型语言模型(LLMs)推理能力的关键技术,但对策略内数据的需求严重限制了训练效率。通过重要性采样(IS)重用策略外数据可以提高效率,但会引入相当大的不稳定性。因此,PPO和GRPO等算法广泛采用IS比率裁剪来稳定训练。然而,训练稳定性和梯度估计主要取决于IS比率与优势的乘积。为进一步稳定训练,我们提出ACPO,它裁剪IS比率与优势的乘积,从而产生更稳定的梯度估计。我们还在ACPO与策略镜像下降(PMD)中的梯度裁剪之间建立了联系,后者是稳定优化过程的标准技术,并证明了在标准RL设置下裁剪PMD的收敛性。在广泛使用的数学推理基准上的实验表明,ACPO在准确性和训练效率方面均持续优于PPO和GRPO,在标准数学基准上使用Qwen3-8B+PPO时带来了4-6个百分点的提升。因此,ACPO是LLMs的RL后训练中传统IS比率裁剪的一种实用且有效的替代方案。

英文摘要

Reinforcement learning (RL) has become a cornerstone for improving the reasoning capabilities of large language models (LLMs), but the need for on-policy data substantially limits training efficiency. Reusing off-policy data through importance sampling (IS) can improve efficiency but introduce considerable instability. Hence, algorithms such as PPO and GRPO widely adopt IS-ratio clipping to stabilize training. However, training stability and gradient estimate are mainly determined by the product of IS ratio and advantage. To further stabilize training, we propose ACPO, which clips the product of the IS ratio and the advantage, leading to more stable gradient estimates. We also establish a connection between ACPO and gradient clipping in policy mirror descent (PMD), which is a standard technique to stabilize optimization process, and prove the convergence of clipped-PMD under the standard RL setting. Experiments on widely used mathematical reasoning benchmarks show that ACPO consistently outperforms PPO and GRPO in both accuracy and training efficiency, delivering 4-6 percentage points gains on standard math benchmarks, with Qwen3-8B+PPO. Hence, ACPO is a practical and effective alternative to conventional IS-ratio clipping for RL post-training of LLMs.

发表机构

  • MIT(麻省理工学院)
  • Shenzhen Loop Area Institute(深圳河套学院)

机构由 AI 辅助整理,请以论文原文为准。

↑