学习使用工具:用于工具集成数学推理的强化学习
Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对Countdown任务,构建工具使用的SFT数据集并采用多种强化学习方法,结合计算器工具集成提升LLM数学推理,使pass@1达66.0%,显著降低计算与验证错误。
AI中文摘要:
当前大型语言模型(LLMs)从外部工具集成中日益受益,尤其是在需要可靠计算与验证的任务中。受此启发,本研究针对倒计时(Countdown)任务,探索计算器工具调用对提升数学推理能力的作用。我们首先分析推理失败案例,发现计算错误占错误响应的很大比例。随后构建监督微调(SFT)数据集,以教会模型实用的工具使用模式及如何解释工具返回的输出。基于此工具格式化策略,我们采用多种在线策略强化学习方法,包括RLOO、RLOO++、GRPO和DAPO,使用可自动验证的最终答案奖励进行训练。为实现更可靠的评估,我们构建了一个全新的包含1024个问题的 held-out Countdown基准,与训练数据无精确重叠。结果表明,计算器工具集成可持续提升SFT及强化学习基线的性能,在pass@k指标上实现约10个百分点的提升。在所有强化学习方法中,Tool-DAPO性能最优,将pass@1从Tool-SFT的35.8%提升至66.0%。进一步分析显示,即便仅提供最终答案奖励,强化学习仍能鼓励更有效的工具使用。这些发现表明,工具集成可减少算术与验证错误,而强化学习则能提升正确推理路径的概率。
英文摘要:
Current large language models (LLMs) increasingly benefit from external tool integration, especially for tasks requiring reliable computation and verification. Motivated by this, we study calculator tool calling for improving mathematical reasoning on the Countdown task. We first analyze reasoning failures and find that calculation errors account for a substantial portion of incorrect responses. We then construct supervised fine-tuning datasets to teach the model useful tool-use patterns and how to interpret returned outputs. Building on this tool-formatted policy, we apply several on-policy reinforcement learning methods, including RLOO, RLOO++, GRPO, and DAPO, using automatically verifiable final-answer rewards. To enable a more reliable evaluation, we construct a fresh 1,024-problem held-out Countdown benchmark with no exact overlap with the training data. Our results show that calculator tool integration consistently improves both SFT and RL baselines, yielding roughly 10 percentage-point gains across pass@k. Among the RL methods, Tool-DAPO achieves the strongest performance, improving pass@1 from 35.8% for Tool-SFT to 66.0%. Further analysis shows that RL encourages more effective tool use even when only final-answer rewards are provided. These findings suggest that tool integration reduces arithmetic and verification errors, while RL increases the probability of correct reasoning traces.