发表机构
University of Tübingen; Tübingen AI Center; Helmholtz Munich; Munich Center for Machine Learning (MCML); Thomson Reuters Foundational Research; Imperial College London; Technical University of Munich(图宾根大学; 图宾根人工智能中心; 亥姆霍兹慕尼黑中心; 慕尼黑机器学习中心; 汤森路透基础研究; 伦敦帝国理工学院; 慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出按需重放(RoD)方法,根据模型学习动态动态分配重放样本,以平衡持续预训练中的适应与遗忘,无需预设重放比例,达到或优于固定重放基线和模型合并。
AI 中文摘要
持续预训练使语言模型能够适应新的领域和知识,但往往以遗忘先前获得的能力为代价。重放可以缓解这种权衡,但固定的重放混合分配训练时独立于模型的实际保留需求。我们引入了按需重放(RoD),它从模型的学习动态中推导出重放分配。RoD通过剩余学习潜力优先处理适应样本,并通过观察到的遗忘优先处理重放样本。它们对共享训练预算的竞争产生了一个在线课程,决定每一步训练什么。跨模型、规模和适应领域,RoD达到或改善了调整后的固定重放基线和模型合并的适应-遗忘前沿,而无需预先规定重放分配。重放集中在更容易遗忘的来源上,并在训练过程中随着遗忘的出现动态增加和重新分配。总之,我们的结果表明,重放可以从模型的演化状态在线分配,针对需要的内容,在需要的时候进行。
英文摘要
Continued pretraining enables language models to adapt to new domains and knowledge, but often at the cost of forgetting previously acquired capabilities. Replay can mitigate this trade-off, but fixed replay mixtures allocate training independently of the model's actual retention needs. We introduce Replay on Demand (RoD), which instead derives the replay allocation from the model's learning dynamics. RoD jointly prioritizes adaptation samples by their remaining learning potential and replay samples by their observed forgetting. Their competition for a shared training budget yields an online curriculum that determines what to train on at each step. Across models, scales, and adaptation domains, RoD reaches or improves upon the adaptation-forgetting frontier of tuned fixed-replay baselines and model merging without prescribing a replay allocation in advance. Replay concentrates on sources that are more vulnerable to forgetting and dynamically increases and redistributes as forgetting emerges during training. Together, our results show that replay can be allocated online from the model's evolving state, targeting what is needed, when it is needed.