arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不覆盖式学习:持续推理中自蒸馏与监督微调的理论

Learning without Overwriting: A Theory of Self-Distillation and Supervised Fine-Tuning in Continual Reasoning

Shinichi Uemura, Taiji Suzuki

arXiv 2610.05200首次发表:更新:

发表机构

Center for Advanced Intelligence Project, RIKEN; Department of Mathematics Informatics, The University of Tokyo(理化学研究所先进智能项目中心; 东京大学数学信息学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文理论分析持续推理中在线策略自蒸馏与监督微调,证明OPSD通过稀疏更新避免遗忘,而SFT密集更新导致灾难性遗忘,并强调预训练多样性的关键作用。

AI 中文摘要

大语言模型(LLM)的在线策略自蒸馏(OPSD)已被证明能够在保留已有知识的同时提升推理能力。尽管有大量实证成功,OPSD在持续推理中的动态机制仍未完全被理解。我们将LLM推理建模为在有向无环图上的搜索,对后训练(持续学习中的OPSD和监督微调SFT)的动态以及预训练对后续性能的影响提供了统一的理论分析。我们的发现确立了三个关键见解,并带有优化保证:(i)带有正确输出提示的OPSD通过提示结构引发的稀疏而有效的梯度下降更新,实现了无遗忘的持续学习。(ii)在正确推理路径上的SFT可能由于沿训练路径的密集更新而导致灾难性遗忘,这些更新覆盖了先前获取的信息。(iii)预训练中的多样性对于使后训练模型在从中间状态开始展开时能够达到正确输出至关重要。我们的结果,由理论分析支持,表明可靠的持续推理取决于后训练更新如何与预训练期间建立的推理结构相互作用。

英文摘要

On-policy self-distillation (OPSD) of large language models (LLMs) has demonstrated the ability to improve reasoning capabilities while preserving previously acquired knowledge. Despite substantial empirical success, the dynamics of OPSD in continual reasoning remain incompletely understood. Modeling LLM reasoning as search over a directed acyclic graph, we provide a unified theoretical analysis of both the dynamics of post-training---OPSD and supervised fine-tuning (SFT) in continual learning---and the impact of pre-training on subsequent performance. Our findings establish three key insights with an optimization guarantee: (i) OPSD with hints from correct outputs enables continual learning without forgetting by sparse yet effective gradient descent updates induced by the hint structure. (ii) SFT on correct reasoning paths can lead to catastrophic forgetting due to dense updates along the training paths, which overwrite the information previously acquired. (iii) Diversity in pre-training is crucial for enabling a post-trained model to reach a correct output when a rollout starts from an intermediate state. Our results, supported by theoretical analysis, show that reliable continual reasoning depends on how post-training updates interact with the reasoning structure established during pre-training.

Commentsmain 11 pages, total 55 pages, main 2 figures, total 4 figures, 1 algorithm table in appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑