arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Latent-LoRA:用于持续学习的具有无梯度路由的紧凑潜在空间适配器

Latent-LoRA: Compact Latent-Space Adapters with Gradient-Free Routing for Continual Learning

Reza Rahimi Azghan, Gautham Krishna Gudur, Giulia Pedrielli, Pavan Turaga, Hassan Ghasemzadeh

arXiv 2607.23837首次发表:更新:

发表机构

College of Health Solutions, Arizona State University; Department of Electrical and Computer Engineering, The University of Texas at Austin; School of Computing and Augmented Intelligence, Arizona State University; The GAME School, Arizona State University(亚利桑那州立大学健康解决方案学院; 德克萨斯大学奥斯汀分校电气与计算机工程系; 亚利桑那州立大学计算与增强智能学院; 亚利桑那州立大学GAME学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型持续学习的灾难性遗忘问题,提出Latent-LoRA系统。利用冻结嵌入层池化令牌嵌入分离任务分布,无需可训练路由组件,通过奇异值分解和正交正则化控制任务干扰,实验显示其有近零遗忘的最优性能。

AI 中文摘要

大型语言模型在单个任务上泛化良好,但缺乏顺序学习的内在机制,导致灾难性遗忘。为缓解此问题,基于LoRA的持续学习方法为每个任务分配单独的低秩适配器,但现有方法要么在推理时需要任务标识,要么不加区分地对所有适配器求和。基于门控的解决方案将输入路由到正确的适配器,但引入了需要防止遗忘的可训练参数。本文观察到,来自冻结的语言模型嵌入层的池化令牌嵌入在整个学习序列中已经分离了任务分布。在这些嵌入上拟合高斯混合模型,无需基于梯度的训练,就足以在测试时进行与任务无关的适配器选择,从而无需学习门控模块。在适配器方面,通过奇异值分解将每个任务的参数约束到预训练权重的主子空间,得到紧凑的潜在空间参数化。在这个子空间内,正交正则化直接控制任务间干扰。由此产生的Latent-LoRA系统无需重放,不需要可训练的路由组件,每个任务使用的参数大大减少。在五个模型规模和两个既定的持续学习基准上的实验表明,该系统具有近零遗忘的最优性能。

英文摘要

Large language models generalize well to individual tasks but lack an inherent mechanism for learning them sequentially, leading to catastrophic forgetting. To mitigate this, LoRA-based continual learning methods allocate a separate low-rank adapter per task, yet existing approaches either require task identity at inference or sum all adapters indiscriminately, letting irrelevant branches distort the output. Recent gating-based solutions route inputs to the correct adapter but introduce trainable parameters that themselves need protection against forgetting. In this work, we observe that pooled token embeddings from a frozen LLM embedding layer already separate task distributions throughout the learning sequence. A Gaussian mixture model fitted on these embeddings, without any gradient-based training, is sufficient for task-agnostic adapter selection at test time. This eliminates the need for a learned gating module. On the adapter side, constraining each task's parameters to the principal subspace of the pretrained weights via SVD yields a compact latent-space parameterization. Within this subspace, orthogonal regularization directly controls inter-task interference. The resulting system, Latent-LoRA, is replay-free, requires no trainable routing component, and uses substantially fewer parameters per task. Experiments across five model scales and two established continual learning benchmarks show state-of-the-art performance with near-zero forgetting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑