arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16850cs.CL

组熵控制策略优化

Group Entropy-Controlled Policy Optimization

Guangran Cheng, Chengqi Lyu, Songyang Gao, Wenwei Zhang, Kai Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对大语言模型强化学习中异构任务的探索利用问题,提出组熵控制策略优化(GEPO)方法,利用组熵进行优势塑造,通过实验验证该方法优于GRPO等,能实现跨任务平衡改进并保持任务探索水平。

中文摘要 AI 辅助

熵控制已成为大语言模型强化学习中的有效工具,有助于在对齐过程中平衡探索与利用的权衡。这种强化学习范式通常在异构任务混合上进行,同一策略下会产生不同熵状态,使得全局或令牌级熵调节不足以满足相应的异构探索需求。这种异质性还使GRPO式归一化优势产生与熵相关的偏差,导致不同提示组的优势信号在统计上不可比。为解决此问题,我们提出组熵控制策略优化(GEPO),这是对GRPO的轻量级扩展,利用从现有分组样本估计的组熵进行熵条件不对称优势塑造。GEPO减弱低熵组的正优势以减少过度利用,减弱高熵组的负优势以保留探索,通过历史熵统计得出自适应阈值。在跨越数学、物理、科学、代码生成和指令跟随的13个基准上对两个基础模型进行的广泛实验表明,GEPO始终优于GRPO和最近的熵控制方法,在整个训练过程中实现了平衡的跨任务改进,同时保持特定任务的探索水平。

英文摘要

Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment process. Such RL paradigm is often conducted on mixtures of heterogeneous tasks, which induce distinct entropy regimes under the same policy, making global or token-level entropy regulation insufficient to corresponding heterogeneous needs of exploration. This heterogeneity further makes GRPO-style normalized advantages induce an entropy-dependent bias, making advantage signals across prompt groups statistically non-comparable. To address this issue, we propose Group Entropy-Controlled Policy Optimization (GEPO), a lightweight extension to GRPO that uses group entropy, estimated from existing grouped samples to perform entropy-conditioned asymmetric advantage shaping. GEPO attenuates positive advantages in low-entropy groups to reduce over-exploitation, and negative advantages in high-entropy groups to preserve exploration, with adaptive thresholds derived from historical entropy statistics. Extensive experiments on two base models across thirteen benchmarks spanning mathematics, physics, science, code generation, and instruction following show that GEPO consistently outperforms GRPO and recent entropy-controlled methods, delivering balanced cross-task improvements while preserving task-specific exploration levels throughout training.

补充信息

↑