发表机构
Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM强化学习后训练中能力遗忘问题,提出CoKL正则化框架,可在目标任务改进与原有能力保留间实现更优平衡,优于现有正则化方法。
AI 中文摘要
强化学习(RL)已成为大语言模型(LLM)后训练的核心范式,但针对新目标的优化会降低基础模型已具备的能力。KL正则化被广泛用于通过约束策略向参考模型漂移来缓解此类遗忘。然而,标准的全策略KL正则化会约束整个响应分布,可能不必要地限制探索和目标任务学习。这引发了一个自然问题:是否可以通过更精确的约束,在保留现有能力的同时最小化对新任务学习的干扰?为此,我们提出了正确性条件KL正则化(CoKL),这是一种将保留约束从完整输出分布缩小到正确性条件响应分布的条件正则化框架。我们用前向KL散度实例化CoKL,并推导了用于基于RL的LLM后训练的实用有限群训练目标。在总体层面,CoKL将分配给正确响应的总概率与其正确性条件分布解耦,从而在不直接锚定错误输出或总正确性质量的情况下,对参考支持的正确响应之间的相对概率分配进行正则化。我们进一步表明,当参考策略不完善时,全策略前向和反向KL正则化会产生严格的最优正确性差距,而CoKL可避免此限制。在受控多解决方案环境和跨多种模型规模的持续后训练设置中的实验表明,与现有正则化方法相比,CoKL在目标任务改进和先前能力保留之间实现了更有利的平衡。我们的代码可在该https URL获取。
英文摘要
Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already present in the base model. KL regularization is widely used to mitigate such forgetting by constraining policy drift toward a reference model. However, standard full-policy KL regularization constrains the entire response distribution and may unnecessarily restrict exploration and target-task learning. This raises a natural question: can a more precise constraint preserve existing capabilities while minimizing interference with learning new tasks? To this end, we propose \underline{Co}rrectness-Conditioned \underline{KL} Regularization (CoKL), a conditional regularization framework that narrows the preservation constraint from the full output distribution to correctness-conditioned response distributions. We instantiate CoKL with forward KL divergence and derive a practical finite-group training objective for RL-based LLM post-training. At the population level, CoKL decouples the total probability assigned to correct responses from their correctness-conditioned distribution, thereby regularizing the relative probability allocation among reference-supported correct responses without directly anchoring incorrect outputs or total correctness mass. We further show that full-policy forward and reverse KL regularization induce a strict optimal correctness gap when the reference policy is imperfect, whereas CoKL avoids this limitation. Experiments in controlled multi-solution environments and continual post-training settings across multiple model scales demonstrate that CoKL achieves a more favorable balance between target-task improvement and prior-capability retention than existing regularization methods. Our code is available at https://github.com/Lumina04/CoKL.