arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08963cs.LGcs.AIcs.CL

关于KL正则化策略优化的研究

On KL-Regularized Policy Optimization

Yifan Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

提出KLPO框架,通过将KL正则化锚定于采样器并利用最小二乘拟合,实现无需评论家、单次rollout的LLM策略优化,统一了SPPO等现有方法。

中文摘要 AI 辅助

异步强化学习(RL)用于大型语言模型(LLM)智能体时,会在由另一个策略生成的轨迹上训练一个策略:轨迹来自过时的检查点,并且即使在相同参数下,推理引擎的概率也与训练器的概率不同。标准的补救措施要么裁剪重要性比率,这会使更新产生偏差,要么像GRPO那样,在每个提示上采样一组响应,这在情节较长时成本高昂。我们提出了KL正则化策略优化(KLPO),一种将KL正则化器锚定在采样器上的框架。正则化的改进步骤随后具有闭式吉布斯解,并且KLPO通过最小二乘法在采样器自身的轨迹上拟合其对数比率最优条件,因此采样器概率通过对数比率进入,无需重要性权重。对回归截距进行剖析,用信号的采样器均值加上采样器到训练器的KL散度替代了难以处理的对数配分函数。对于词元级策略镜像下降目标,我们表明即使存在随机工具输出,所得梯度也可以仅从终端回报计算,无需评论家,通过以采样器为中心的分数或单条轨迹残差。我们进一步证明,KL项的独立蒙特卡洛估计使这些梯度无偏,推导了更廉价的top-$K$和二元近似的精确KL差距,并表明SPPO、GPO、REBEL和BPO是KLPO的特例。结果是一种无需评论家的更新,每个提示仅使用一次 rollout,既不需要学习归一化器,也不需要一组响应。

英文摘要

Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters. Standard remedies either clip importance ratios, which biases the update, or, as in GRPO, sample a group of responses per prompt, which is costly when episodes are long. We propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler. The regularized improvement step then has a closed-form Gibbs solution, and KLPO fits its log-ratio optimality condition by least squares on the sampler's own trajectories, so the sampler probability enters through a log-ratio and no importance weights are needed. Profiling out the regression intercept replaces the intractable log-partition function with the signal's sampler mean plus a sampler-to-trainer KL divergence. For token-level policy mirror descent targets, we show that the resulting gradient can be computed from terminal returns without a critic, via sampler-centered scores or a single trajectory residual, even under stochastic tool outputs. We further prove that independent Monte Carlo estimates of the KL term keep these gradients unbiased, derive the exact KL gap of cheaper top-$K$ and binary approximations, and show that SPPO, GPO, REBEL, and BPO arise as special cases of KLPO. The result is a critic-free update that uses one rollout per prompt and requires neither a learned normalizer nor a group of responses.

发表机构

  • Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑