arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EvoThink:通过自剪枝和顿悟时刻偏好优化在大型推理模型中进化思维

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization

Xinbang Dai, Zheyu Xin, Huikang Hu, Lin Ren, Rihui Jin, Guohui Xiao, Guilin Qi, Kuicai Dong, Zhaocheng Du, Yuyang Zhang

arXiv 2607.19962首次发表:更新:

发表机构

Southeast University; Noah’s Ark Lab(东南大学; 诺亚方舟实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大型推理模型过度思考问题,提出EvoThink框架,含自剪枝训练和顿悟时刻偏好优化两个关键组件,能减少冗余验证、探索新推理路径,经评估可提高推理效率与能力。

AI 中文摘要

大型推理模型(LRMs)常因冗余验证步骤而过度思考。现有缓解过度思考的方法,如快慢思维切换和推理轨迹压缩,无法在LRM推理过程中对有益和冗余步骤进行细粒度区分,可能损害推理能力。为同时提高推理效率和能力,我们提出EvoThink框架,它减少冗余验证并鼓励探索新推理路径。EvoThink包括两个关键组件:自剪枝训练(SPT),一种无监督方法,迭代修剪冗余推理步骤并在简洁轨迹上自训练;顿悟时刻偏好优化(AMPO),受遗传算法启发,识别有价值的失败推理尝试,合成从错误到正确的顿悟时刻数据,并优化模型以内化此推理模式。在数学推理和代码生成基准上的广泛评估表明,EvoThink不仅大幅减少推理时的令牌使用,还提高了LRMs的推理能力。

英文摘要

Large Reasoning Models (LRMs) often suffer from overthinking due to redundant verification steps. Existing approaches for mitigating overthinking, such as fast-slow thinking switching and reasoning trajectory compression, fail to make a fine-grained distinction between beneficial and redundant steps within the LRM's reasoning process, and may thus impair reasoning capability in their pursuit of efficiency. To simultaneously improve reasoning efficiency and capability, we propose EvoThink, a framework that reduces redundant verification and encourages the exploration of new reasoning paths. EvoThink comprises two key components: Self-Pruning Training (SPT), an unsupervised method that iteratively prunes redundant reasoning steps and self-trains on the concise trajectories; and Aha-Moment Preference Optimization (AMPO), which, inspired by genetic algorithms, identifies valuable failed reasoning attempts, synthesizes from-wrong-to-right aha-moment data, and optimizes the model to internalize this reasoning pattern. Extensive evaluations across mathematical reasoning and code generation benchmarks demonstrate that EvoThink not only substantially reduces inference-time token usage but also improves the reasoning capability of LRMs.

Comments9 pages, 7 figures, accepted by IJCAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑