惊人成功,反复失败:熵引导的信用分配用于LLM推理中的探索
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning
浏览论文内容
中文总结 AI 辅助
针对LLM推理中RLVR信用分配问题,提出熵引导的EAPO方法,不对称处理成功与失败,利用熵信号强化惊人成功并纠正反复失败,无需额外监督,在推理任务上取得最佳性能并提升探索多样性。
中文摘要 AI 辅助
基于可验证奖励的强化学习(RLVR)通过结果级反馈增强了大语言模型(LLM)的推理能力,然而近期更细粒度的信用分配方法往往需要辅助模型、额外采样或特权信息。尽管策略熵提供了现成的信号,但在强化和惩罚下优先处理不确定位置会将惩罚集中在失败响应仍保留恢复替代方案的位置,这可能抑制探索机会。为解决此问题,我们引入了熵优势策略优化(EAPO),一种将成功与失败不对称对待的熵引导信用分配方法。具体而言,受“不确定性下的成功较难重复,而自信的失败往往重现”这一观察的启发,EAPO将归一化策略熵与响应优势的符号相结合,以强化惊人成功并纠正反复失败。它对成功响应中的高熵决策分配更强的强化,对失败响应中的低熵决策分配更强的惩罚,同时减弱不确定位置的惩罚以保留恢复机会。通过跨令牌重新分配响应优势,EAPO直接从现有 rollout 信号中推导出令牌级信用,无需额外监督。我们在基础和推理骨干网络上的一系列推理任务上验证了EAPO,证明其取得了最佳整体性能。我们进一步表明,EAPO促进了更有效的探索,拓宽了问题覆盖范围并生成了更多样化的候选答案。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) enhances reasoning in large language models (LLMs) through outcome-level feedback, yet recent approaches to finer-grained credit assignment often require auxiliary models, additional sampling, or privileged information. Although policy entropy provides a readily available signal, prioritizing uncertain positions under both reinforcement and penalization concentrates penalties where failed responses still retain alternatives for recovery, which can suppress opportunities for exploration. To address this, we introduce Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method that treats success and failure asymmetrically. Specifically, motivated by the observation that success under uncertainty is less repeatable while confident failures tend to recur, EAPO couples normalized policy entropy with the sign of the response advantage to reinforce surprising success and correct repeated failure. It assigns stronger reinforcement to high-entropy decisions in successful responses and stronger penalties to low-entropy decisions in failed responses, while attenuating penalties at uncertain positions to preserve opportunities for recovery. By redistributing the response advantage across tokens, EAPO derives token-level credit directly from existing rollout signals without additional supervision. We validate EAPO on a range of reasoning tasks across both base and reasoning backbones, demonstrating that it achieves the best overall performance. We further show that EAPO promotes more effective exploration, broadening problem coverage and generating more diverse candidate answers.
发表机构
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。