GVPO++:面向大语言模型后训练与同策略蒸馏的组方差策略优化
GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation
浏览论文内容
中文总结 AI 辅助
针对LLM后训练中重要性采样导致的不稳定问题,提出GVPO方法,通过KL约束奖励最大化解耦梯度权重,保证唯一最优解并支持灵活采样,且可扩展至同策略蒸馏。
中文摘要 AI 辅助
后训练在增强大型语言模型(LLMs)的推理能力和任务特定专长方面发挥着关键作用。尽管后训练方法(如组相对策略优化(GRPO))近期取得了进展,但其实际部署仍因依赖重要性采样而导致的训练不稳定性而受到阻碍。我们引入了组方差策略优化(GVPO),一种新颖的后训练方法,它将KL约束奖励最大化的解析解整合到其梯度加权方案中。该公式提供了一个直观的解释:GVPO的梯度对应于隐式奖励的中心距离与实际奖励的中心距离之间的均方误差。GVPO提供两个关键优势:(1)它保证唯一的最优解,精确对应于KL约束的奖励最大化目标;(2)它支持灵活的采样分布,无需重要性采样。超越一般后训练,我们展示了GVPO自然扩展到同策略蒸馏(OPD)。此外,GVPO能够优化广泛的扩展OPD目标族,为多样化的目标设计提供了原则性基础。通过统一理论保证与实际适应性,GVPO为可靠且多功能的LLM后训练和同策略蒸馏建立了新范式。
英文摘要
Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Despite recent advances in post-training methods, such as Group Relative Policy Optimization (GRPO), their practical deployment remains impeded by training instability arising from the reliance on importance sampling. We introduce Group Variance Policy Optimization (GVPO), a novel post-training method that integrates the analytical solution of KL-constrained reward maximization into its gradient weighting scheme. This formulation provides an intuitive interpretation: GVPO's gradient corresponds to the mean squared error between the central distance of implicit rewards and that of actual rewards. GVPO offers two key advantages: (1) it guarantees a unique optimal solution, exactly to the KL-constrained reward maximization objective, and (2) it enables flexible sampling distributions without requiring importance sampling. Beyond general post-training, we show that GVPO naturally extends to on-policy distillation (OPD). Furthermore, GVPO enables the optimization of a broad family of extended OPD objectives, providing a principled foundation for diverse objective design. By unifying theoretical guarantees with practical adaptability, GVPO establishes a new paradigm for reliable and versatile LLM post-training and on-policy distillation.