语言模型能否在其权重中持续学习事实?
Can a Language Model Learn Facts Continually in Its Weights?
浏览论文内容
中文总结 AI 辅助
研究语言模型能否在权重中持续学习事实,通过追踪写入Qwen3模型的虚构事实及实验发现,训练数据广度决定知识类型,事实可行为遗忘但不擦除,广泛数据能创建可用知识,可靠通道是上下文而非权重。
中文摘要 AI 辅助
持续学习有望使语言模型在训练后不断获取知识,将每个新事实写入其权重中。权重写入是否能支持知识积累仍未确定。我们追踪了从创建到后续二十至一百次写入序列中写入Qwen3模型的虚构事实,使用五种类型的预留问题,以在提示中给出事实的原始模型作为参考。在这些实验中,训练数据的广度决定了所创建知识的类型。简单陈述训练产生背诵,而多样的重述将背诵与应用的差距从27.4分降至5.4分,且未向模型展示结论。这种差异延续到后续写入:经过二十次连续写入后,简单陈述事实的准确率保持在1%,而从广泛研究数据写入的事实准确率保持在46%。我们还发现事实可以在行为上被遗忘而不被擦除。被遗忘的事实保留了其写入时添加的大部分对数概率,在简单陈述训练下,70%关于它们的错误答案包含最近写入的事实。相同的写入几乎不会降低模型在上下文中对事实的应用,提示中提供的被遗忘的研究事实在其问题上的准确率恢复到77 - 80%。这些结果描述了存储但由问题键控的知识:后续写入会重定向到达它的问题。对不相关能力的损害与原始模型的KL散度相关,后续写入会产生干扰,无论早期事实如何存储。广泛的数据可以创建可用的知识,冻结的参考可以保留能力,但我们测试的任何干预措施,包括基于每次写入的准确局部测量构建的措施,都无法使早期事实可访问。当事实必须被组合或在后续写入中幸存时,可靠的通道是上下文而不是权重。
英文摘要
Continual learning is a long-standing capability gap between LLMs and humans. Writing new knowledge into a model's weights routinely causes it to forget old knowledge, commonly denoted as "catastrophic forgetting". Various modifications of supervised fine-tuning and distillation aim to mitigate catastrophic forgetting, but quantifying what (or how much) information was forgotten is often difficult. In this paper, we study whether current methods of writing knowledge into weights enable models to learn continually without forgetting. We introduce a framework for studying continual learning in the iterative regime, writing invented facts one at a time into a Qwen3 model already modified by previous writes, and varying the training data, method, and parameter update. Across SFT and off- and on-policy distillation, using LoRA or full fine-tuning, we compare repeated statement training (the same fact repeated in two formats) with varied example training (24 factual restatements) and find that varied examples comprehensively support more flexible use. After twenty sequential writes and merges, the model answers only 1% of questions about earlier facts correctly when every write uses repeated statements, compared with 46% when every write uses varied examples. We additionally show that this retention depends on the data used for the later writes, regardless of training method or parameter update, and that behavioral forgetting of an earlier fact does not erase its presence from the log-probabilities. Together, our framework neatly provides a comparison of performance across training data, training regimes, and parameter update schemes in an iterative learning task.
发表机构
- Baseten(巴斯滕)
机构由 AI 辅助整理,请以论文原文为准。