arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

局部支持学习

Local Support Learning

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

arXiv 2610.02126首次发表:更新:

发表机构

MIT CSAIL; Tel Aviv University(麻省理工学院计算机科学与人工智能实验室; 特拉维夫大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出局部支持学习(LSL),通过门控机制使梯度更新局部于当前数据分布,解决大规模预训练模型(高达70亿参数)在多个训练阶段中的灾难性遗忘,同时高效保留预训练与微调能力。

AI 中文摘要

我们在大规模预训练模型的背景下探索灾难性遗忘问题。通过将遗忘视为每个权重矩阵输入空间中的几何问题,我们揭示了一种自然的保留目标,在该目标下,基于梯度的优化器产生的更新是次优的。基于这一观察,我们提出了局部支持学习(LSL),这是一个通用框架,用于增强基于梯度的训练,以在无法访问先前数据的情况下保留先前能力。在新的学习阶段,LSL 将两个具有不同角色的组件配对:一个标准的权重适配器,照常训练以最小化损失,以及一个门控函数,该函数仅在其自身训练分布的输入激活上启用适配器,使更新局部于该分布。关键挑战在于,该门必须在仅使用当前阶段数据训练的同时,路由来自所有学习阶段的数据。我们通过基于高斯混合模型(GMM)的门来解决这一问题,其似然在其训练数据附近迅速衰减,从而自然地倾向于在来自先前阶段的数据上保持关闭。我们表明,这种训练后方法可以解决高达 70 亿参数的 LLM 中的遗忘问题,在多个训练阶段中保留预训练和微调能力,同时在内存和计算方面高效,对超参数选择稳健,并展现出扩展潜力。

英文摘要

We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge is that this gate must route data from all learning phases while training only on data from the current one. We address this with a gate based on a Gaussian Mixture Model (GMM), whose likelihood decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We show that this post-training approach can resolve forgetting in LLMs of up to 7 billion parameters, retaining both pretrained and finetuned capabilities across multiple training phases, while being efficient in memory and compute, robust to hyperparameter choice, and showing scaling potential.

CommentsWebsite and code: https://assafbk.github.io/lsl

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑