反思性恢复:一种通过从错误中学习进行推理的自监督方法
Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes
浏览论文内容
中文总结 AI 辅助
针对LLM推理中模仿学习因错误累积导致的规模塌缩问题,提出反思性恢复自监督方法,利用失败轨迹片段训练模型纠错,显著提升基准性能并实现自我纠正。
中文摘要 AI 辅助
数据驱动的微调因其简单高效而被广泛用于增强大型语言模型(LLMs)的推理能力。然而,主流模仿学习方法仅依赖完美的推理轨迹,会遭遇“规模塌缩”问题:当问题集有限时,增加正例无法带来持续的性能提升。此外,在推理过程中,LLM无法保证每一步中间步骤都是正确的,因此容易出错。一旦出现此类错误,LLM往往难以恢复,并可能因先前错误的累积而进一步被误导。为解决这一问题,我们提出了反思性恢复(Reflective Recovery),一种简单而有效的自监督方法,将失败的推理尝试转化为恢复训练数据。具体来说,我们提取失败轨迹的初始片段,将其与提示词拼接,并用以引导LLM走向有效解决方案。由于这些来自失败轨迹的片段很可能包含错误,该过程教会模型在推理过程中识别并纠正错误,从而无需依赖外部批评者或奖励模型即可从错误状态中恢复。在广泛基准上的评估表明,反思性恢复显著提升了性能。在DeepSeek-R1-Distill-Qwen-7B上,它在AIME 2025上的准确率从30.0%提升至37.5%,在Minerva上从37.6%提升至47.8%。更重要的是,分析表明它打破了规模塌缩的障碍,使模型能够发展出涌现性的自我纠正行为,代表了从面向结果的记忆到面向过程的反思性推理的范式转变。
英文摘要
Data-driven fine-tuning is widely adopted to enhance reasoning in Large Language Models (LLMs) due to its simplicity and efficiency. However, mainstream imitation learning methods that rely exclusively on perfect reasoning trajectories suffer from a Scaling Collapse: when the problem set is limited, increasing positive examples fails to yield continuous improvement. However, during inference, an LLM can not guarantee that every intermediate step is correct and is therefore prone to errors. Once such errors arise, the LLM often struggles to recover and may be further misled by the accumulation of previous mistakes. To address this, we propose Reflective Recovery, a simple yet effective self-supervised approach that transforms failed reasoning attempts into recovery training data. Specifically, we extract initial segments of failed trajectories, concatenate them with prompts, and use them to guide the LLM toward valid solutions. Because these segments from failed trajectories are likely to contain errors, this process teaches models to recognize and correct mistakes during reasoning, enabling recovery from erroneous states without relying on external critics or reward models. Evaluated on extensive benchmarks, Reflective Recovery significantly improves performance. On DeepSeek-R1-Distill-Qwen-7B, it boosts accuracy from 30.0% to 37.5% on AIME 2025 and from 37.6% to 47.8% on Minerva. More importantly, analyses demonstrate that it breaks the scaling collapse barrier and enables models to develop emergent self-correction behaviors, representing a paradigm shift from outcome-oriented memorization to process-oriented reflective reasoning.
发表机构
- Zhejiang University(浙江大学)
- The University of Hong Kong(香港大学)
- The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。