AI 中文总结
本研究提出SCLoRA方法,通过结合频谱裁剪与低秩适配(LoRA),在提升下游任务性能的同时缓解了灾难性遗忘,实现更好的学习且减少知识遗忘。
AI 中文摘要
近年来,低秩适配(LoRA)已成为一种重要范式,该范式冻结预训练权重,引入小型可学习适配器,而非微调全部参数。本研究基于奇异值分解(SVD)揭示了网络参数奇异分量的若干关键见解:其一,预训练网络参数中具有大奇异值的主奇异分量可在微调期间有效复用,而具有较小奇异值的次要分量更具任务特异性,需要大量适配;其二,我们首次建立理论联系,即LoRA适配器中奇异值的不受控增长会导致预训练知识的遗忘,这是一种被称为灾难性遗忘的知名问题。基于这些观察,我们提出SCLoRA,该方法以感知预训练模型频谱分布的方式,将带参数化的奇异分量与频谱裁剪注入预训练模型。SCLoRA通过将更新聚焦于需要适配的分量,有效适配新任务,同时缓解灾难性遗忘。我们开展了大量实验,结果表明SCLoRA不仅提升了下游任务性能,还能有效保留预训练知识。
英文摘要
In recent years, low-rank adaptation (LoRA) has emerged as a significant paradigm that freezes pre-trained weights and introduces small, learnable adapters instead of fine-tuning the full set of parameters. In this work, we uncover several key insights regarding the singular components of network parameters based on Singular Value Decomposition (SVD). Firstly, the principal singular components with large singular values in pre-trained network parameters can be effectively reused during fine-tuning, whereas the minor components with smaller singular values are more task-specific and require substantial adaptation. Secondly, we first establish the theoretical connection that the uncontrolled growth of singular values in LoRA adapters leads to the forgetting of pre-trained knowledge -- a well-known issue referred to as catastrophic forgetting. Building on these observations, we propose SCLoRA, which injects parameterized singular components with spectral clipping into the pre-trained model in a way that is aware of the spectral distribution of the pre-trained model. SCLoRA effectively adapts to new tasks by focusing updates on components that require adaptation, while simultaneously alleviating catastrophic forgetting. We conduct extensive experiments and demonstrate that SCLoRA not only improves downstream performance but also effectively retains pre-trained knowledge.
CommentsACL 2026 Main Conference