当KL正则化在组策略优化中失效时
When KL Regularization Misfires in Group Policy Optimization
浏览论文内容
中文总结 AI 辅助
针对组策略优化中KL正则化失效的问题,分析七种失效模式,提出ZCPO方法,经数学推理实验和消融研究验证其有效性。
中文摘要 AI 辅助
为何移除参考策略KL正则化有时会提升组策略优化效果?这促使我们研究参考策略信息应如何融入组相对更新。我们分析了KL与奖励相互作用中的七种潜在失效模式:奖励裁剪后的残余KL更新、梯度抵消后的残余KL更新、奖励相同的组中的KL更新;KL随响应长度增长及其相对贡献的不平衡;KL集中在少量token上;以及将k1纳入奖励时的采样噪声。我们提出零和校准策略优化(ZCPO),其使用条件KL测量的相对漂移来校准组内奖励系数,并将其整合到基础替代函数中。数学推理实验和消融研究支持该设计在我们设定下的有效性。
英文摘要
Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into the base surrogate. Mathematical reasoning experiments and ablations support this design's effectiveness in our settings.
发表机构
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。