arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从结果到策略:学习数学推理的策略效用

From Outcomes to Strategies: Learning Strategy Utility for Mathematical Reasoning

Ruikang Zhang, Xiao An, Xuli Shen, Jiaxing Sun, Xiaoyi Yu, Jin Zeng, Jiang Wu, Tong Lin

arXiv 2609.32482首次发表:更新:

AI 中文总结

本文提出SURE框架,通过分离高级策略与推理过程,学习策略效用以指导强化学习,在三个策略主干上平均提升pass@1达2.93%,并降低计算成本。

AI 中文摘要

使用可验证奖励的强化学习已显著提升了数学推理能力。然而,仅凭最终正确性,在将高级策略(如定理选择和子目标分解)与其后续执行分开考虑时,对策略质量提供的洞察有限。本文研究策略效用,其定义为在给定执行器下,策略支持正确下游解决方案的可能性。我们引入SURE,一个用于学习和利用相对策略效用的框架。在该框架中,高级策略与其详细推理过程相分离。基于从策略条件化轨迹和教师生成的对比中构建的成对偏好,学习一个策略奖励模型来估计相对策略效用。在强化学习期间,冻结的奖励模型仅读取提取的策略,其得分与正确性和格式奖励在序列级GRPO目标中相结合。与仅基于结果和格式的GRPO基线相比,实验表明SURE在三个策略主干上的平均pass@1分别提高了1.87%、2.64%和2.93%。我们的方法在准确率上也达到或优于更强的奖励基线,同时所需的GRPO阶段计算量显著更低。

英文摘要

Reinforcement learning with verifiable rewards has substantially improved mathematical reasoning. However, terminal correctness alone provides limited insight into the quality of high-level strategies, such as theorem selection and subgoal decomposition, when considered separately from their subsequent execution. This paper studies strategy utility, which is defined as the likelihood that a strategy supports a correct downstream solution under a given executor. We introduce SURE, a framework for learning and leveraging relative strategy utility. In this framework, high-level strategies are separated from their detailed reasoning. Based on the pairwise preferences constructed from strategy-conditioned rollouts and teacher-generated contrasts, a Strategy Reward Model is learned to estimate relative strategy utility. During reinforcement learning, the frozen reward model reads only the extracted strategy, whose score is combined with the correctness and format rewards in a sequence-level GRPO objective. Compared with outcome-and-format GRPO baselines, experiments show that SURE improves average pass@1 by 1.87%, 2.64%, and 2.93% across three policy backbones. Our method also achieves competitive or better accuracy than stronger reward baselines while requiring substantially lower GRPO-stage compute.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑