AI 中文总结
本文提出CPLUS方法,利用冻结模型的自探针生成数据并记录预测,以梯度信号而非重放数据来缩减冲突更新,在五个语言模型和四个基准上显著减少遗忘,且随模型规模增大效果更佳。
AI 中文摘要
将预训练模型适应新数据可能导致对先前学习行为的灾难性遗忘。当仅剩少量过去样本时,它们为持续学习方法提供了关于保留内容的稀疏且狭窄的证据。我们表明,语言模型可以通过自探针(self-probing)扩展这一证据,其中冻结模型从保留样本生成新输入,并记录其对这些输入的自身预测。与先前将此类数据作为训练示例进行重放的工作不同,我们的方法CPLUS利用自探针和过去样本梯度来缩减与先前行为冲突的参数更新。在四个基准上使用五个语言模型的实验显示了三个结果。首先,相同的探针作为梯度信号比作为重放数据更能保留先前行为。其次,CPLUS在学习新数据的同时,始终比现有基线减少更多遗忘,尤其是在过去数据稀缺时,并且这种保护扩展到未用于训练的基准。第三,我们观察到CPLUS随着模型增大而变得更加有效:在Qwen3模型家族内,它恢复了标准微调所导致遗忘的越来越多的份额。
英文摘要
Adapting pretrained models to new data can cause catastrophic forgetting of previously learned behavior. When only a few past samples remain, they give continual learning methods sparse and narrow evidence about what to preserve. We show that language models can expand this evidence through self-probing, in which the frozen model generates new inputs from the retained samples and records its own predictions on them. Unlike prior work that replays such data as training examples, our method, CPLUS uses self-probe and past-sample gradients to scale down parameter updates that conflict with prior behavior. Experiments with five language models on four benchmarks show three results. First, the same probes preserve more prior behavior as gradient signals than as replay data. Second, CPLUS learns the new data while consistently reducing forgetting more than existing baselines, especially when past data are scarce, and this protection extends to benchmarks not used for training. Third, we observe that CPLUS also becomes more effective as models grow: within the Qwen3 model family, it recovers an increasing share of the forgetting caused by standard fine-tuning.
Comments27 pages, 18 figures