arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

震耳欲聋的沉默:灾难性遗忘存在于数据从未提及的令牌的输出嵌入中

A Deafening Silence: Catastrophic Forgetting Lives in the Output Embeddings of Tokens the Data Never Speaks

Jonghyun Han, Younghoon Song, Jongyoul Park

arXiv 2610.09835首次发表:更新:

发表机构

Seoul National University of Science and Technology; Korea Institute of Land & Infrastructure Safety Technology(首尔科学技术大学; 韩国国土与基础设施安全技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示LLM灾难性遗忘集中于罕见令牌的输出嵌入,提出仅提高输出投影Adam epsilon的干预,在多种设置下消除39.4%-67.9%遗忘,且不损害目标学习。

AI 中文摘要

大型语言模型(LLMs)的持续预训练和微调不可避免地引发灾难性遗忘,通常通过使用往往无法访问的原始数据进行重放来缓解。在这种无数据场景下,我们分析了遗忘发生的位置及其原因。在多达1.4B参数的五个设置中进行系统性参数冻结,揭示遗忘集中选择性地发生在新语料库中罕见出现的令牌的输出嵌入中,而主体的相同sqrt(v-hat)频带是惰性的,新学习驻留在其他位置。这种定位由语料库的词汇不足决定,而非训练模式,从而允许在固定基础模型内仅根据令牌计数进行预训练风险排序。机制上,缺失令牌接收持续的单侧softmax梯度,Adam的二阶矩(sqrt(v-hat))归一化将其放大为全尺寸更新。因此,我们提出一种干预措施:在训练期间仅提高输出投影的Adam epsilon。在跨越160M到12B参数和四个模型家族的八个设置中,这消除了所有七个稳定配置中39.4%至67.9%的遗忘,而不损害目标学习或需要逐模型调优。该防御与重放相加或更好地结合(在Qwen/韩语上达到79.8%),并拯救了发布的头部LoRA免受23倍遗忘激增。由于对漂移行的事后编辑恢复不到5%的遗忘,干预必须在训练期间进行。我们的发现表明,单行优化器调整可能作为对抗灾难性遗忘的主要防御,其中语料库使词汇匮乏。

英文摘要

Continual pre-training and fine-tuning in Large Language Models (LLMs) inevitably induce catastrophic forgetting, typically mitigated by replay using often-inaccessible original data. In this data-free regime, we analyze where forgetting occurs and why. Systematic parameter freezing across five settings up to 1.4B reveals that forgetting concentrates selectively in the output embeddings of tokens rarely seen in the new corpus, whereas the same sqrt(v-hat) band of the body is inert and new learning resides elsewhere. This localization is governed by the vocabulary deficiency of the corpus rather than the training mode, allowing pre-retraining risk ranking from token counts alone within a fixed base model. Mechanistically, absent tokens receive persistent one-sided softmax gradients that Adam's second-moment (sqrt(v-hat)) normalization amplifies into full-sized updates. We therefore propose an intervention: raising Adam's epsilon exclusively for the output projection during training. Across eight settings spanning 160M to 12B parameters and four model families, this removes 39.4% to 67.9% of forgetting across all seven stable configurations without degrading target learning or requiring per-model tuning. The defense combines additively or better with replay (79.8% on Qwen/Korean) and rescues released-head LoRA from a 23-fold forgetting surge. Because post-hoc editing of the drifted rows recovers under 5% of forgetting, the intervention must operate during training. Our findings indicate that a single-line optimizer adjustment may serve as the primary defense against catastrophic forgetting where the corpus starves the vocabulary.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑