PaLoRA:面向持续学习的分步低秩自适应
PaLoRA: Paced Low-Rank Adaptation for Continual Learning
浏览论文内容
中文总结 AI 辅助
PaLoRA通过推导步进律$s^*=\sqrt{R/c}$,提出秩感知的自适应梯度缩放与SVD截断及零空间投影相结合,在持续学习中平衡稳定性与可塑性,在50任务基准上提升4%准确率。
中文摘要 AI 辅助
基于LoRA的持续学习方法通过各种机制缓解灾难性遗忘,但几乎所有方法都辅以较小的学习率作为启发式手段来限制梯度缩放幅度。这种固定启发式缺乏关于限制强度应如何随任务累积而演化的理论指导。我们揭示,即使在零空间投影等方向性约束下,有限精度更新也不可避免地沿多个方向泄漏到累积先验知识的子空间中。虽然较小的学习率能减弱此类泄漏,但随着历史知识的有效秩增长,它们无法阻止累积遗忘的加剧。我们表明,最优的幅度限制应随该有效秩自适应地增加,以平衡稳定性与可塑性,即对先前知识的保留与新任务信息的获取。在各向异性泄漏模型下,我们推导出步进律$s^*=\sqrt{R/c}$,该定律刻画了梯度步长的最优缩放,即幅度限制本身,其中$R$是过去更新的有效秩。基于这一见解,我们提出PaLoRA,它通过自适应SVD截断压缩历史知识,将梯度投影到先前任务的零空间,并应用秩感知的自适应步进。实验表明,与先前方法相比,该方法取得了一致的改进,在长时程设置中表现尤为突出,在具有挑战性的50任务ImageNet-A和ImageNet-R基准上实现了4%准确率的显著提升。
英文摘要
LoRA-based continual learning methods mitigate catastrophic forgetting through various mechanisms, yet nearly all complement these with small learning rates as a heuristic to restrict gradient scaling magnitude. Such fixed heuristics lack theoretical guidance on how the strength of this restriction should evolve as tasks accumulate. We reveal that even under directional constraints such as nullspace projection, finite-precision updates inevitably leak into the subspace of accumulated prior knowledge along multiple directions. While small learning rates attenuate such leakage, they cannot prevent the accumulated forgetting from intensifying as the effective rank of historical knowledge grows. We show that the optimal magnitude restriction should adaptively increase with this effective rank to balance stability and plasticity, i.e., preservation of previous knowledge and acquisition of new task information. Under an anisotropic leakage model, we derive a pacing law $s^*=\sqrt{R/c}$ that characterizes the optimal scaling of gradient steps, i.e., the magnitude restriction itself, where $R$ is the effective rank of past updates. Based on this insight, we propose PaLoRA, which compresses historical knowledge via adaptive SVD truncation, projects gradients onto the nullspace of prior tasks, and applies rank-aware adaptive pacing. Experiments demonstrate consistent improvements over prior methods, with particularly strong performance in long-horizon settings, achieving substantial gains of 4% accuracy on challenging 50-task ImageNet-A and ImageNet-R benchmarks.
发表机构
- Peking University(北京大学)
- Chinese Academy of Sciences(中国科学院)
机构由 AI 辅助整理,请以论文原文为准。