arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型的时间增量持续预训练:无灾难性遗忘的知识更新

Time-Incremental Continued Pretraining of LLMs: Knowledge Updates Without Catastrophic Forgetting

Fırat Öncel, Salman Hussain Ali, Mirco Ravanelli, Cem Subakan, Çağatay Yıldız

arXiv 2609.23916首次发表:更新:

发表机构

Concordia University; Mila – Quebec AI Institute; Université de Montréal; Laval University; University of Tübingen; Tübingen AI Center(康考迪亚大学; 米拉-魁北克人工智能研究所; 蒙特利尔大学; 拉瓦尔大学; 蒂宾根大学; 蒂宾根人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在真实网络爬取数据上评估时间增量持续预训练,发现其能有效获取知识且无灾难性遗忘,数据质量优于数量,LoRA可匹配完整CPT,且增益可迁移至下游任务。

AI 中文摘要

大语言模型(LLMs)在其预训练结束的那一刻便逐渐过时,然而从头开始重新训练的成本高得令人望而却步。持续预训练(CPT)是自然的补救措施,但通常通过持续学习的视角进行评估,该视角假设数据流是不相交的。这对于网络规模爬取数据上的时间增量更新来说并不适用,因为连续快照在设计上共享大量URL重叠。我们在这一现实场景中研究时间增量CPT:在FineWeb-Edu数据转储上进行持续预训练,这些数据严格来自每个模型知识截止日期之后,评估涵盖三个系列(OLMo2、Llama-3.1/3.2、Gemma-3-1B)的六个开放权重模型和四个参数规模(1B-3B-7B-8B)。我们的发现围绕四个实际问题展开。(i)是否获得了知识?是的,但具有异质性,且没有灾难性遗忘:六个模型中有五个在截止日期前的事实回忆上也有所改善,且增益与预训练饱和度相关(主要由每参数令牌预算驱动)。(ii)成本是多少?几乎为零:在十三项任务套件上的宏观平均值对每个模型都保持在基线的0.01以内。(iii)配方是什么?数据质量主导数量(一个精选的60亿令牌切片与更广泛的400亿令牌切片相匹配);知识获取和通用能力的最优学习率相差约一个数量级;且足够秩的LoRA与完整CPT相匹配。(iv)它能经受部署吗?CPT增益通过SFT转移,而DPO的效果则因系列而异。总之,这些结果描绘了比先前持续学习文献所建议的更乐观的时间增量CPT图景。

英文摘要

Large language models (LLMs) drift out of date the moment their pretraining ends, yet retraining from scratch is prohibitively expensive. Continued pretraining (CPT) is the natural remedy, but it is typically evaluated through a continual learning lens that assumes disjoint data streams. This is a poor fit for time-incremental updates on web-scale crawls, where successive snapshots share substantial URL overlap by design. We study time-incremental CPT in this realistic regime: continued pretraining on FineWeb-Edu dumps drawn strictly from after each model's knowledge cutoff, evaluated across six open-weight models spanning three families (OLMo2, Llama-3.1/3.2, Gemma-3-1B) and four parameter scales (1B-3B-7B-8B). We organize our findings around four practical questions. (i) Is knowledge acquired? Yes, but heterogeneously, and without catastrophic forgetting: five of six models also improve on pre-cutoff factual recall, and the gains track pretraining saturation (driven primarily by token budget per parameter). (ii) What does it cost? Almost nothing: the macro-average across a thirteen-task suite stays within 0.01 of the base for every model. (iii) What is the recipe? Data quality dominates quantity (a curated 6B-token slice matches a broader 40B one); the optima for knowledge acquisition and general capability are separated by roughly an order of magnitude in learning rate; and LoRA at sufficient rank matches full CPT. (iv) Does it survive deployment? CPT gains transfer through SFT, while DPO's effect is family-dependent. Together, these results paint a more optimistic picture of time-incremental CPT than the prior continual learning literature suggests.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑