arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33181cs.AIcs.CL

SeOPD:通过自生成思维链的在线策略蒸馏实现大语言模型自我进化

SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought

Xiaoshu Chen, Sihang Zhou, Ke Liang, Xinwang Liu

AI总结:

本文提出SeOPD,利用单个大语言模型自身的深度思考模式生成思维链作为特权信息,通过在线策略蒸馏指导非思考模式,实现无需外部信息的自我进化,提升模型整体推理能力。

AI中文摘要:

近期在线策略自我蒸馏(OPSD)的进展表明,大语言模型(LLMs)可以通过利用外部特权信息(PI)(如人工标注或外部环境的反馈)来提升自身能力。然而,获取准确的标注和构建复杂的环境通常需要大量的人力和计算资源,这限制了OPSD的可扩展性。尽管一些近期研究探索了在没有外部PI的情况下进行自我改进,但所取得的收益仍然有限。在这项工作中,我们探讨大语言模型是否能在没有外部PI的情况下实现相当的自我改进。我们的关键观察是,单个大语言模型可以支持多种推理模式,例如深度思考模式和非思考模式,其中深度思考在推理过程中会生成额外信息。基于这一观察,我们提出了自我进化在线策略蒸馏(SeOPD),使大语言模型能够蒸馏并内化由自身思维链(CoT)生成的信息。具体而言,它(1)使用深度思考模式生成思维链,(2)使用非思考模式生成响应,以及(3)将生成的思维链作为PI,为非思考响应提供词元级监督,使得推理过程中推断出的新信息能够引导非思考模式并内化到共享模型参数中,从而同时提升非思考和深度思考能力。跨大语言模型和任务的大量实验证明了SeOPD的有效性。

英文摘要:

Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without external PI, the resulting gains remain limited. In this work, we explore whether LLMs can achieve comparable self-improvement without external PI. Our key observation is that a single LLM can support multiple reasoning modes, such as deep-thinking and non-thinking modes, with deep thinking generating additional information during reasoning. Based on this observation, we propose Self-Evolving Online Policy Distillation (SeOPD), which enables LLMs to distill and internalize information generated by their own chain of thought (CoT). Specifically, it (1) generates CoT with the deep-thinking mode, (2) produces responses with the non-thinking mode, and (3) uses the generated CoT as PI to provide token-level supervision for the non-thinking response, allowing new information inferred during reasoning to guide the non-thinking mode and be internalized into the shared model parameters, thereby improving both non-thinking and deep-thinking capabilities. Extensive experiments across LLMs and tasks demonstrate the effectiveness of SeOPD.

↑