arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05872cs.LGcs.AIcs.CL

离线策略合并优于在线策略自蒸馏用于持续学习

Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning

Chen Henry Wu, Thomas Zhang, Aditi Raghunathan

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出“嫁接”方法,通过离线策略合并优于在线策略自蒸馏,在持续学习中实现新任务与旧任务性能的帕累托改进,并避免昂贵采样。

中文摘要 AI 辅助

人工智能的一个长期目标是构建能够持续学习并自我改进的模型。在预训练后的模型上,对新数据进行监督微调(SFT)往往导致泛化能力差和灾难性遗忘。因此,传统观点认为在线策略训练是持续学习的先决条件。然而,在实践中,包含新知识或能力的数据往往是离线策略的。虽然诸如在线策略自蒸馏(OPSD)等方法试图通过将离线策略数据转换为在线策略信号来弥合这一差距,但已有研究表明这些方法会导致推理崩溃。在本文中,我们证明离线策略合并优于OPSD用于持续学习。我们首先表明SFT能从新数据中学习到有用的信号,但简单地应用其更新会干扰现有能力。我们通过一种称为“嫁接”(grafting)的简单方法减少这种干扰,该方法改变了更新的学习位置和应用方式:(1)在较早的捐赠者检查点上学习更新,理想情况下甚至在预训练结束之前,并将权重更新应用于预训练后的模型;(2)缩放权重更新,这等效于一种模型合并形式;(3)可选地,当新数据分布远离预训练后的模型时,屏蔽最敏感的更新方向。在多种持续学习设置中,包括(1)从专家轨迹中蒸馏,(2)使用STaR和教学强化学习进行自我改进,以及(3)在预训练截止后注入知识,嫁接方法在新任务和旧任务性能上均帕累托优于SFT和OPSD,同时避免了昂贵的在线策略采样。因此,我们的工作挑战了在线策略训练作为强化学习训练模型持续学习的必要条件的观点。

英文摘要

A long-standing goal of AI is a model that can continually learn and improve itself. On post-trained models, supervised finetuning (SFT) on new data often causes poor generalization and catastrophic forgetting. As such, the conventional wisdom is that on-policy training is a prerequisite for continual learning. In practice, however, data containing new knowledge or capabilities are often off-policy. While methods such as on-policy self-distillation (OPSD) try to bridge this gap by converting off-policy data into on-policy signal, they have been shown to cause reasoning collapse. In this paper, we show that off-policy merging beats OPSD for continual learning. We first show that SFT learns a useful signal from new data, but naively applying its update interferes with existing capabilities. We reduce this interference with a simple recipe we term grafting, which changes where the update is learned and how it is applied: (1) learning the update on an earlier donor checkpoint, ideally even before the end of pretraining, and applying the weight update to the post-trained model; (2) scaling the weight update, equivalent to a form of model merging; and (3) optionally, masking the most sensitive update directions when the new data distribution is far from the post-trained model. Across continual learning settings including (1) distilling from expert traces, (2) self-improvement with STaR and Pedagogical RL, and (3) injecting knowledge after pretraining cutoff, grafting Pareto-dominates both SFT and OPSD in new-task and old-task performance, while avoiding expensive on-policy sampling. Therefore, our work challenges on-policy training as a necessity for continual learning on RL-trained models.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

↑