发表机构
National University of Singapore; The Hong Kong University of Science and Technology; University of Science and Technology of China(新加坡国立大学; 香港科技大学; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于概率奖励的无验证器强化学习在长程推理中因后验集中现象导致的奖励不可区分问题,提出RLCPR框架,通过不确定性感知采样和集中感知正则化提升优化稳定性与令牌效率,在多数基准上超越基线。
AI 中文摘要
基于概率奖励的无验证器强化学习为在缺乏外部验证器的一般推理任务上训练大型语言模型提供了一种有前景的方法。然而,这些奖励的可靠性,尤其是在长程推理中的可靠性,仍未得到充分探索。本研究识别出概率奖励的一种长度依赖的失效模式,我们称之为后验集中现象(PCP)。我们表明,当推理轨迹变长时,参考答案在给定推理轨迹条件下的概率往往坍缩到一个低方差区间。这一现象导致奖励几乎无法区分,在基于GRPO的设置下,使得基于概率的策略优化变得不稳定且低效。受此启发,我们提出了具有集中感知后验奖励的强化学习(RLCPR),这是一个无验证器的强化学习框架,明确考虑PCP以提升优化稳定性和令牌效率。它包含两个组件:不确定性感知的数据采样,在生成前减少易发生集中的轨迹;以及集中感知的正则化,在后验奖励坍缩时惩罚不必要地长的轨迹。大量实验表明,除了更高的令牌效率外,RLCPR在七个基准中的六个上(包括通用领域和数学推理挑战)比最先进的无验证器强化学习基线高出最多4.0%。
英文摘要
Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by this, we propose Reinforcement Learning with Concentration-aware Posterior Rewards (RLCPR), a verifier-free RL framework to explicitly account for PCP for better optimization stability and token efficiency. It has two components: uncertainty-aware data sampling, which reduces concentration-prone rollouts before generation, and concentration-aware regularization, which penalizes unnecessarily long traces when posterior rewards collapse. Extensive experiments show that, alongside higher token efficiency, RLCPR outperforms the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks, including general-domain and mathematical reasoning challenges.
CommentsFix typo