arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12161cs.LGcs.CL

当KL正则化在组策略优化中失效时

When KL Regularization Misfires in Group Policy Optimization

Fei Ding

首次发表
浏览论文内容

中文总结 AI 辅助

针对组策略优化中KL正则化失效的问题,分析七种失效模式,提出ZCPO方法,经数学推理实验和消融研究验证其有效性。

中文摘要 AI 辅助

为何移除参考策略KL正则化有时会提升组策略优化效果?这促使我们研究参考策略信息应如何融入组相对更新。我们分析了KL与奖励相互作用中的七种潜在失效模式:奖励裁剪后的残余KL更新、梯度抵消后的残余KL更新、奖励相同的组中的KL更新;KL随响应长度增长及其相对贡献的不平衡;KL集中在少量token上;以及将k1纳入奖励时的采样噪声。我们提出零和校准策略优化(ZCPO),其使用条件KL测量的相对漂移来校准组内奖励系数,并将其整合到基础替代函数中。数学推理实验和消融研究支持该设计在我们设定下的有效性。

英文摘要

Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into the base surrogate. Mathematical reasoning experiments and ablations support this design's effectiveness in our settings.

发表机构

  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑