发表机构
School of artificial intelligence, Wuhan University; School of computer science, Shanghai Jiao Tong University; Institute of automation, Chinese Academy of Sciences(武汉大学人工智能学院; 上海交通大学计算机科学与工程学院; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出 SRPO 框架,将人类的自反思能力内化到 LLMs 中,无需额外模型即可将稀疏监督转化为密集训练信号,在数学推理等基准上以高数据效率取得了最先进性能。
AI 中文摘要
自反思是人类学习中用于信用分配的强大机制,能将稀疏的结果反馈转化为可操作的指导。然而,其在训练后的大语言模型(LLMs)中的潜力仍未得到充分探索。我们提出自反思策略优化(SRPO)框架,该框架将此能力内化。SRPO使LLMs能够分析自身完成的轨迹,将错误综合为简洁的“反思补丁”,并利用基于反思的教师对学生在线 rollout 的评分作为密集的 token 级训练信号。此过程无需外部评论家、单独的奖励模型或更大的教师模型,即可有效将稀疏的终端监督转化为密集的 token 级学习信号。我们在数学推理和长程智能体基准测试中证明,SRPO达到了最先进的性能,且数据效率极高。使用 Qwen3-8B 基础模型,SRPO仅用扩展监督微调所需训练 FLOPs 的 8%(0.08 倍),就在 AIME'24 上达到 73.3%,同时显著提高了 WebShop(64.7%)、ALFWorld(76.8%)和 SWE-Bench-Lite(31.2%)的成功率。代码可在此 https URL 获取。
英文摘要
Self-reflection is a powerful mechanism for credit assignment in human learning, converting sparse outcome feedback into actionable guidance. However, its potential for post-training Large Language Models (LLMs) remains underexplored. We propose Self-Reflective Policy Optimization (SRPO), a framework that internalizes this capability. SRPO enables LLMs to analyze their own completed trajectories, synthesize errors into concise "reflection patches," and use reflection-conditioned teacher scores on student on-policy rollouts as dense token-level training signals. This process effectively transforms sparse terminal supervision into dense, token-level learning signals without requiring external critics, separate reward models, or larger teacher models. We demonstrate that SRPO achieves state-of-the-art performance across mathematical reasoning and long-horizon agentic benchmarks with exceptional data efficiency. Using a Qwen3-8B base model, SRPO attains 73.3% on AIME'24 using only 8% (0.08x) of the training FLOPs required by scaled supervised fine-tuning, while significantly improving success rates on WebShop (64.7%), ALFWorld (76.8%), and SWE-Bench-Lite (31.2%). Code is available at https://github.com/Galleons2029/SRPO
CommentsAccepted to ICML 2026