发表机构
Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有 LLM 推理强化学习方法依赖稀疏结果奖励的局限,提出 ConsensusBench 基准及 ConsensusPR 算法,通过共识节点提供过程级奖励,在多推理数据集上优于 GRPO 类方法。
AI 中文摘要
强化学习(RL)已成为提升大语言模型(LLM)推理能力的主要范式之一,其中组相对策略优化(GRPO)及相关算法凭借结果级奖励展现出强劲性能。然而,这些方法仅依赖最终答案,无法获取哪些中间步骤促成成功或失败的反馈。随着任务复杂度和推理轨迹长度增加,这种稀疏的最终答案奖励愈发不足。为解决此局限,我们推出 ConsensusBench——一个旨在提供基于规则的过程级信号的新型数据集。我们假设正确的最终答案依赖于推理过程中的一小部分中间结论,这些结论可视为可验证的子结果。我们通过从 N 次 rollout 中筛选正确轨迹并对语义等价的中间陈述进行聚类来识别这些子结果,将这些聚类后的陈述称为共识节点(Consensus Nodes)。通过将源自这些节点的基于规则的过程奖励集成到 GRPO 类算法中,我们开发出名为 ConsensusPR 的新型强化学习信号,它直接降低了长推理轨迹中结果奖励的稀疏性。为便于系统的过程级评估,我们在基准中引入三个指标:最终答案准确率(Acc)、节点覆盖率(NCR)及每个节点的标记数(TPN)。在 AIME 2024、AIME 2025、GSM8K、MATH-500 及我们的 ConsensusBench 上开展的实验表明,所提方法始终优于 GRPO 类方法,凸显了共识节点在指导推理方面的实用价值。
英文摘要
Reinforcement learning (RL) has become one of the primary paradigms for reasoning enhancement of large language models (LLMs). In particular, Group Relative Policy Optimization (GRPO) and related algorithms have demonstrated strong performance with outcome-level rewards. However, these methods depend solely on the final answer, without feedback regarding which intermediate steps contribute to success or failure. As task complexity and reasoning trajectory length increase, such sparse final-answer rewards become increasingly insufficient. To address this limitation, we introduce ConsensusBench, a novel dataset designed to provide rule-based process-level signals. We posit that a correct final answer relies on a small set of intermediate conclusions throughout the reasoning process, which can be seen as a verifiable sub-outcome. We identify these sub-outcomes by filtering correct trajectories from N rollouts and clustering semantically equivalent intermediate statements. We call these clustered statements as Consensus Nodes. By integrating a rule-based process reward derived from these nodes into GRPO-style algorithms, we develop a new reinforcement learning signal named ConsensusPR. It directly reduces the reward sparsity of outcome reward across long reasoning trajectories. To facilitate systematic process-level evaluation, we introduce three metrics to our benchmark: Final Answer Accuracy (Acc), Node Coverage Rate (NCR), and Tokens per Node (TPN). Experiments across AIME 2024, AIME 2025, GSM8K, MATH-500, and our ConsensusBench demonstrate that the proposed method consistently surpasses GRPO-style approaches, highlighting the practical value of consensus nodes in guiding reasoning.