发表机构
State Key Laboratory of Novel Software Technology, Nanjing University; Shanghai Artificial Intelligence Laboratory(南京大学计算机软件新技术国家重点实验室; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对DAPO中动态采样的训练效率问题,提出DAA方法整合得到DA3PO,大幅优于GRPO及其他GRPO变体
AI 中文摘要
解耦的CLIP与动态采样策略优化(DAPO)是分组相对策略优化(GRPO)的重要变体,相比GRPO有多项改进,其中动态采样对DAPO准确率提升的贡献最大。为提升准确率,动态采样通过消除零优势导致的零策略梯度来增强训练稳定性,具体是过滤掉采样响应完全正确或完全错误的提示,避免这类零梯度。但我们的理论分析显示,动态采样会降低训练效率,因为它无法有效利用困难提示上的难采样正确响应,形式上会非对称放大同一提示不同响应的优势:困难提示上错误响应的放大程度大于正确响应,导致模型回避生成观测到的错误响应,而非利用困难提示上的难采样正确响应,进而造成训练效率低下。为提升训练效率,我们提出直接优势放大(DAA),用于放大动态采样得到的困难提示上难采样正确响应的优势,确保使用动态采样时能有效利用这些难采样响应,提高训练效率。将DAA整合到DAPO后,我们得到难度感知优势放大策略优化(DA3PO),其基于DAPO实现,代码不足30行。实验表明,DA3PO的性能显著优于GRPO及其他经典GRPO变体。
英文摘要
Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) is a prominent variant of Group Relative Policy Optimization (GRPO). DAPO introduces several improvements over GRPO. Among these, Dynamic Sampling contributes the most to DAPO's accuracy gains relative to GRPO. To improve accuracy, Dynamic Sampling enhances training stability by eliminating zero policy gradients from zero advantages. Specifically, it avoids such zero gradients by filtering out prompts where sampled responses are either entirely correct or incorrect. However, our theoretical analysis shows that Dynamic Sampling decrease training efficiency as it cannot effectively utilize hard-to-sample correct responses on hard prompts. Formally, it asymmetrically amplifies the advantages of distinct responses to the same prompts. On hard prompts, incorrect responses undergo greater amplification than correct ones. This leads the model to avoid generating the observed incorrect responses rather than capitalizing on the hard-to-sample correct ones on hard prompts, resulting in low training efficiency. To improve training efficiency, we propose Direct Advantage Amplification (DAA), which amplifies the advantages of hard-to-sample correct responses on hard prompts, as obtained by Dynamic Sampling. This ensures that, when Dynamic Sampling is used, these hard-to-sample responses can be effectively capitalized on, implying higher training efficiency. By integrating DAA into DAPO, we obtain Difficulty-aware Advantage Amplification Policy Optimization (DA3PO), which is implemented with fewer than 30 lines of code from DAPO. Experiments show that DA3PO significantly outperforms GRPO and other classical GRPO variants.