arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02139cs.CLcs.AIcs.LG

基于渐进式经验演化的自我改进大语言模型

Self-Improving Large Language Models via Progressive Experience Evolution

Shijie Ren, Xiting Wang, Meng Li, Yujie Guo, Yunhang Yao, Ziheng Peng, Xunlong Wang, Yuetan Chen, Haoyang Zhou, Yunlong Liang, Fandong Meng

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出统一后训练框架SPEE,通过显式经验演化与隐式策略优化结合,在五种数学推理基准上提升了大语言模型的自我改进能力,优于现有基线方法。

中文摘要 AI 辅助

具备自我改进能力的大语言模型(LLMs)不仅需要有效的策略优化,还需要一种将瞬时交互经验转化为持久模型能力的原则性机制。现有的自我改进范式仍处于碎片化状态:测试时方法可显式提取经验,但无法将其内化到模型参数中;而训练时优化方法可更新模型参数,但缺乏积累可迁移经验的显式机制。弥合这两种范式需要一个关键的中间阶段,即经验蒸馏,该阶段尚未得到充分探索。为解决这一差距,我们提出了SPEE(Self-Progressive Experience Evolution,自我渐进式经验演化),这是一个统一的后训练框架,依次执行显式经验演化和隐式策略优化。在显式经验演化阶段,SPEE对从多次交互中收集的轨迹进行反思,以提取、验证并渐进式演化可迁移经验,随后通过特权引导的同策略自蒸馏(OPSD)将这些经验内化到策略中。在隐式策略优化阶段,基于奖励的强化学习利用这些内化的先验来探索新的解决方案策略。在经验演化阶段,一个持续演化的全局经验池整合来自成功和失败轨迹的知识,过滤低效用经验,并缓解由单个轨迹引发的事后合理化问题。在五个数学推理基准上的实验表明,SPEE在三种模型规模下均始终优于测试时和训练时的自我演化基线。源代码可在该https URL获取。

英文摘要

Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction experience into persistent model capabilities. Existing self-improvement paradigms remain fragmented: test-time methods can explicitly extract experience but cannot internalize it into model parameters, whereas training-time optimization methods can update model parameters but lack an explicit mechanism for accumulating transferable experience. Bridging these two paradigms requires a critical intermediate stage that remains underexplored, namely \emph{experience distillation}. To address this gap, we propose \textbf{SPEE} (\textbf{S}elf-\textbf{P}rogressive \textbf{E}xperience \textbf{E}volution), a unified post-training framework that sequentially performs explicit experience evolution followed by implicit policy optimization. During explicit experience evolution, SPEE reflects on trajectories collected from multiple interactions to extract, verify, and progressively evolve transferable experience, which is subsequently internalized into the policy through privilege-guided On-Policy Self-Distillation (OPSD). During implicit policy optimization, reward-driven reinforcement learning leverages these internalized priors to explore novel solution strategies. In the experience evolution stage, a continuously evolving global experience pool consolidates knowledge from both successful and failed trajectories, filters out low-utility experience, and mitigates post-hoc rationalization induced by individual trajectories. Experiments on five mathematical reasoning benchmarks demonstrate that SPEE consistently outperforms both test-time and training-time self-evolution baselines across three model scales. The source code is available at https://github.com/rrrsj/SPEE.

补充信息

↑