arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

何时复习:语言模型持续预训练的间隔重复方法

When to Review: Spaced Repetition for Continual Pre-Training of Language Models

Alankar Atreya, Devesh Batra, Yoages Kumar Mantri, Geremy Bantug, Greig A Cowan, Raad Khraishi

arXiv 2608.17530首次发表:更新:

发表机构

University College London; NatWest AI Research(伦敦大学学院; 国民西敏寺银行人工智能研究部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出受认知科学启发的SRT框架,采用SM-2算法调度样本重放,可恢复大型语言模型持续预训练中5至37个百分点的旧知识准确率,同时保留或提升新知识获取能力。

AI 中文摘要

大型语言模型的持续预训练必须在获取新信息的同时不遗忘旧知识。现有重放方法通常选择全局的新旧数据混合比例并均匀采样,却忽略了不同样本的遗忘速度存在差异。我们将持续预训练建模为自适应复习调度问题:训练循环不仅要决定重放的历史数据量,还要确定每一步应重放哪些样本。我们提出间隔重复训练(Spaced Repetition Training,SRT)这一受认知科学启发的持续学习框架,它采用SuperMemo-2(SM-2)算法调度样本重放。SRT维护每个样本的复习状态,将每个样本的困惑度映射为回忆质量信号,调度历史样本以保留知识、调度新样本以巩固知识,同时不改变模型、目标函数和优化器。在时间上分离的维基百科和代码语料库上,SRT改善了稳定性-可塑性权衡,在不同模型规模下,将朴素持续预训练丢失的旧知识准确率恢复了5至37个百分点,同时保留或提升了新知识的获取能力。在更大规模下,SRT保留了朴素持续预训练和均匀重放会大幅下降的广泛基准性能。在视觉和表格数据上的实验进一步表明,当搭配合适的回忆信号时,该调度原理可扩展至语言之外的领域。

英文摘要

Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often choose a global old/new mixture and sample uniformly, ignoring that examples differ in how quickly they are forgotten. We formulate continual pre-training as adaptive review scheduling: the training loop should decide not only how much history to replay, but which examples should return at each step. We introduce Spaced Repetition Training (SRT), a continual learning framework inspired by cognitive science, which schedules sample-rehearsal using the SuperMemo-2 (SM-2) algorithm. SRT maintains per-example review state, maps per-example perplexity to a recall-quality signal, and schedules historical examples for retention and new examples for consolidation while leaving the model, objective, and optimizer unchanged. On temporally separated Wikipedia and code corpora, SRT improves the stability-plasticity trade-off, recovering 5 to 37 percentage points of old-knowledge accuracy lost by naive continual pre-training across model scales while preserving or improving new-knowledge acquisition. At larger scale, SRT preserves broad benchmark performance that naive continual pre-training and uniform replay substantially degrade. Experiments with vision and tabular data further suggest that the scheduling principle extends beyond language when paired with an appropriate recall signal.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑