arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

惊人简单的自我反思无需强化学习即可提升智能体模型

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

Jonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer, Emiliano Penaloza, Zhengyan Shi, Alessandro Sordoni, Marc-Alexandre Côté, Xingdi Yuan, Minseon Kim

arXiv 2609.35741首次发表:更新:

发表机构

RPI; UC San Diego; KAIST; Mila; Microsoft Research(伦斯勒理工学院; 加州大学圣迭戈分校; 韩国科学技术院; 米拉研究所; 微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出仅反思微调(ROFT)方法,仅用智能体自身经历的解释进行训练,无需强化学习即可提升其后续行动能力,在软件工程任务上达到与GRPO相当甚至更优的解决率,并证明学习解释能改善行动。

AI 中文摘要

人们不仅通过重复成功的行为来学习,还通过复述和解释自己的经历来学习,修正自己的理解以指导未来的行为。语言模型智能体能否仅通过训练其自身经历的解释来改进未来的行动?我们通过研究“仅反思微调”(ROFT)来探究这一问题,这是一种极简的在线流程,旨在隔离仅解释训练对后续行为的影响。智能体尝试执行一项任务,观察可获得的反馈,生成一段反思性解释,并仅基于解释的标记以下一词预测损失进行微调。该流程既不使用外部教师,也不使用基于奖励的策略更新。在使用Qwen3.5-4B进行的软件工程实验中,ROFT在混合了基础模型成功与失败尝试的问题上进行训练。在保留的SWE-bench Verified和Pro基准上,经过20次更新后,ROFT在不使用验证器的情况下达到了49.2%和26.8%的解决率,而评估运行中GRPO在40次更新后分别为48.0%和25.3%,并且在训练时间和采样尝试方面取得了更快的早期进展。它还能学会解决所有64次采样的基础模型尝试均失败的任务,表明学习可以在没有任何初始成功轨迹的情况下开始。行为分析发现,ROFT间接地为行动分配信用,鼓励好的行动并抑制错误的行动。此外,提示反思强调更直接的解决方案,即使在无显式长度惩罚的情况下,也会产生更短的后续尝试。总之,这些发现表明,学习解释也能改进学习行动,确立了自我生成的反思作为有用的训练目标,并激发了对解释到行动迁移的进一步研究。

英文摘要

People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.

Comments62 pages, 18 figures, 5 tables, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑