发表机构
University College London; Chinese Academy of Sciences(伦敦大学学院; 中国科学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出函数结构化图强化学习(FSG-RL),结合函数图、Python代码和多验证器反馈,通过GRPO优化策略,在数学推理基准上将最终答案准确率从43.25%提升至67.50%。
AI 中文摘要
算法数学推理需要可靠的分解、计算和聚合。最终答案奖励对中间错误的指导有限,而成功执行并不能保证数学正确性。本工作提出函数结构化图强化学习(FSG-RL),将子问题图与Python实现和多验证器反馈相结合。策略首先通过监督微调(SFT)学习从函数图生成代码。然后,组相对策略优化(GRPO)利用答案门控奖励和跨度级信用分配来优化策略。该框架还支持教师监督和结构化记忆。从Grade School Math 8K(GSM8K)、MathQA、MATH和Omni-MATH中整理出的基准将公共函数图与私有验证规范配对。在统一评估协议下,与SFT相比,GRPO将最终答案准确率从43.25%提高到67.50%,完整解决方案成功率从32.25%提高到52.25%。使用教师监督的持续强化学习(RL)带来了额外的收益。这些收益不仅限于生成格式正确的代码,还支持验证器引导的强化学习用于数学推理。代码可在以下网址获取:此https URL。
英文摘要
Algorithmic mathematical reasoning requires reliable decomposition, computation, and aggregation. Final-answer rewards provide limited guidance on intermediate errors, while successful execution does not guarantee mathematical correctness. This work proposes Function-Structured Graph Reinforcement Learning (FSG-RL), connecting subproblem graphs and Python implementations with multi-verifier feedback. The policy first learns to generate code from function graphs through supervised fine-tuning (SFT). Group Relative Policy Optimization (GRPO) then optimizes the policy using answer-gated rewards and span-level credit assignment. The framework also supports teacher supervision and structured memory. A benchmark curated from Grade School Math 8K (GSM8K), MathQA, MATH, and Omni-MATH pairs public function graphs with private verification specifications. Under a unified evaluation protocol, GRPO improves final-answer accuracy from 43.25% to 67.50% and full solution success from 32.25% to 52.25% over SFT. Continued reinforcement learning (RL) with teacher supervision yields additional gains. The gains extend beyond producing correctly formatted code, supporting verifier-guided reinforcement learning for mathematical reasoning. Code is available at https://github.com/ZihanLiummyycc/FSG-RL.
Comments5 pages, 2 figures, 2 tables