arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越稳定性-探索困境:面向大语言模型策略优化的环境正则化方法

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

Xianlei Zhou, Xiangdi Meng, Yu He, Tianyu Qi, Shuyan Guan, Xianli Zhang, Jian Zhang, Xin Li, Qika Lin, Jun Liu

arXiv 2608.23311首次发表:更新:

发表机构

Alibaba Group; Beijing Normal University; National University of Singapore(阿里巴巴集团; 北京师范大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型策略优化的稳定性-探索困境,提出ERPO方法,通过查询KL散度正则化输入侧分布,接入现有策略优化流水线,在数学推理基准上实现了更强准确率与更稳定行为。

AI 中文摘要

大语言模型(LLM)的策略优化(PO)面临稳定性与探索性之间的权衡,目前通过动作侧的策略KL散度(Policy-KL)正则化器进行调节。这让从业者陷入两难:保留Policy-KL会约束响应行为并消耗动作侧的探索预算,而舍弃它则会让优化失去显式的漂移控制。本文提出一种替代方案,通过将正则化转移到输入侧来打破这一困境。随着训练推进,当前策略诱导的训练查询分布会不受控制地偏离强化学习(RL)前的参考分布。具体而言,环境正则化策略优化(ERPO)引入查询KL散度(QKL)项,用于约束该查询分布偏移,同时采用数据集静态参考派生的逐查询权重,使每个逐查询更新偏向参考下的典型查询。QKL的梯度仅通过查询似然流动;策略梯度估计器使用的响应得分函数未出现在QKL项中,因此QKL不会对响应分布施加直接梯度压力——探索性得以保留。ERPO可接入GRPO/PPO/REINFORCE风格的流水线,无需额外前向传播。在6个数学推理基准上,ERPO替代了标准Policy-KL正则化器,同时实现了对查询分布漂移的有效控制,在高温解码和长视野场景下取得了更强的准确率和更稳定的行为。源代码可在指定链接获取。

英文摘要

Policy optimization (PO) for Large Language Models faces a stability--exploration trade-off, currently mediated by an action-side Policy-KL regularizer. This puts practitioners in a double bind: keeping Policy-KL constrains response behavior and consumes the action-side exploration budget, while dropping it leaves the optimization without an explicit drift control. We argue for an alternative that breaks the dilemma by moving regularization to the input side. As training progresses, the distribution over training queries induced by the current policy drifts unchecked from its pre-RL reference distribution. Concretely, Environment-Regularized Policy Optimization (ERPO) introduces a Query-KL (QKL) term that bounds this query distribution shift, together with a dataset-static reference-derived per-query weight that biases each per-query update toward queries typical under the reference. The QKL gradient flows strictly through the query likelihood; the response score function used by policy-gradient estimators does not appear in the QKL term, so QKL exerts no direct gradient pressure on the response distribution---exploration is preserved. ERPO plugs into GRPO/PPO/REINFORCE-style pipelines without additional forward passes. On six mathematical reasoning benchmarks, ERPO replaces the standard Policy-KL regularizer while achieving effective control over query distribution drift, delivering stronger accuracy and substantially more stable behavior under high-temperature decoding and long-horizon training. Our source code are available at https://github.com/AlibabaResearch/ERPO

CommentsAccepted to EMNLP 2026 main conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑