arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

持续学习中LoRA的动力学理论

A Dynamical Theory of LoRA in Continual Learning

Théo Marchetta, Filippo Alessandroni, Alessandro Breccia, Alessandro Ingrosso, Federica Gerace

arXiv 2609.39367首次发表:更新:

发表机构

Alma Mater Studiorum – Università di Bologna; University College London; Radboud University(博洛尼亚大学; 伦敦大学学院; 拉德堡德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过可解教师-学生模型解析刻画了持续学习中LoRA的动力学,揭示了其减少干扰但减慢适应的双重效应,并提出冻结强第一任务表示的掩蔽策略以显著降低遗忘,同时阐明适配器秩对迁移与遗忘的影响。

AI 中文摘要

尽管低秩适应(LoRA)被广泛使用,但对其在持续学习中的动力学以及低秩更新影响灾难性遗忘的机制知之甚少。我们在一个可解的两任务教师-学生模型中提供了LoRA的渐近精确动力学刻画。在高维在线学习极限下,我们为一组有限的宏观序参量推导出一个封闭的常微分方程组,从而在初始任务1学习阶段和随后任务2上的LoRA微调阶段都得到了泛化误差的精确表达式。该理论定量匹配有限维模拟,并揭示了LoRA的两个特征性效应:低秩适应减少了与第一任务所学特征的干扰,但其初始化减慢了向第二任务的适应。基于这一机制性图景,我们分析了一种状态相关的掩蔽策略,该策略冻结携带最强第一任务表示的隐藏单元,并将适应限制在互补子空间内。这种结构划分显著减少了遗忘,同时保持了新任务上的可塑性。我们的框架进一步阐明了适配器秩的作用:迁移仅在目标任务的内在维度内改善,超过该维度后趋于饱和,而遗忘则随秩增加而持续增长。这些结果为低秩适应如何跨顺序任务组织信息提供了动力学和几何学解释,并在顺序MNIST基准上得到了定性复现。

英文摘要

Despite the widespread use of Low-Rank Adaptation (LoRA), little is known about its dynamics in continual learning and the mechanisms by which low-rank updates affect catastrophic forgetting. We provide an asymptotically exact dynamical characterization of LoRA in a solvable two-task teacher-student model. In the high-dimensional online-learning limit, we derive a closed system of ordinary differential equations for a finite set of macroscopic order parameters, yielding exact expressions for the generalization errors throughout both the initial Task 1 learning phase and the subsequent LoRA fine-tuning on Task 2. The theory quantitatively matches finite-dimensional simulations and exposes two characteristic effects of LoRA: low-rank adaptation reduces interference with features learned on the first task, but its initialization slows adaptation to the second task. Building on this mechanistic picture, we analyze a state-dependent masking strategy that freezes hidden units carrying the strongest first-task representations and restricts adaptation to the complementary subspace. This structural partitioning markedly reduces forgetting, while preserving plasticity on the new task. Our framework further clarifies the role of adapter rank: transfer improves only up to the intrinsic dimensionality of the target task and saturates beyond it, while forgetting continues to grow with rank. These results provide a dynamical and geometric account of how low-rank adaptation organizes information across sequential tasks and are qualitatively reproduced on a sequential MNIST benchmark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑