arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04011cs.AI

探索保持的策略优化

Exploration-Preserving Policy Optimization

Hangzhan jin, Mohammad Hamdaqa, Doina Precup

首次发表
浏览论文内容

中文总结 AI 辅助

针对强化学习中信用分配导致探索不足的问题,提出ExPPO方法,通过基于响应意外度和通过率的优势塑形,提升推理覆盖与正确模式发现,实验验证其有效性。

中文摘要 AI 辅助

具有可验证奖励的强化学习提升了推理能力,而学习信号的分配方式决定了在重复采样下哪些解决方案仍然可及。组相对目标为同等奖励的响应分配同等优势,使得聚合信用与采样模式频率成正比。我们引入了探索保持的策略优化(ExPPO),一种轻量级的优势塑形规则,利用提示相对、长度归一化的响应意外度和提示通过率来重新分配信用。ExPPO将有界塑形与共享归一化相结合,以保持验证器极性并近似维持每个提示组的绝对序列优势总量。我们的分析刻画了响应级信用分配与采样模式更新的关系,推导出熵增益和正确模式发现的局部条件。实验表明,域内和域外推理覆盖均有提升,聚合响应准确率更高,且在大采样预算下覆盖能力强。一项受控的多答案评估进一步展示了正确模式产出增加以及已验证正确响应中多样性的提升。代码可在该 https URL 获取。

英文摘要

Reinforcement learning with verifiable rewards improves reasoning, while the allocation of learning signal shapes which solutions remain accessible under repeated sampling. Group-relative objectives assign equal advantages to equally rewarded responses, making aggregate credit proportional to sampled mode frequency. We introduce Exploration-Preserving Policy Optimization (ExPPO), a lightweight advantage-shaping rule that redistributes credit using prompt-relative, length-normalized response surprisal and prompt pass rate. ExPPO combines bounded shaping with shared normalization to preserve verifier polarity and approximately maintain each prompt group's total absolute sequence-advantage mass. Our analysis characterizes response-level credit allocation alongside sampled mode updates, deriving local conditions for gains in entropy and correct-mode discovery. Experiments show improved in-domain and out-of-domain reasoning coverage, higher aggregate response accuracy, and strong coverage at large sampling budgets. A controlled multi-answer evaluation further demonstrates increased correct-mode yield and gains in diversity among verified-correct responses. Code is available at https://github.com/jinhangzhan/ExPPO

发表机构

  • Mila - Quebec AI Institute(Mila - 魁北克人工智能研究所)
  • Polytechnique Montréal(蒙特利尔理工学院)
  • McGill University(麦吉尔大学)
  • Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑