发表机构
Beijing University of Posts and Telecommunications; National University of Singapore(北京邮电大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型持续学习中的正交性困境,提出CoDe-LoRA方法,通过整合通用知识与解耦任务特定知识,利用零空间投影和语义路由,在多个基准上取得最佳平均准确率。
AI 中文摘要
持续学习(CL)对于大语言模型(LLMs)顺序适应不断演化的任务至关重要。为缓解灾难性遗忘,近期研究采用带正交投影的低秩适配(如O-LoRA)来隔离任务参数。然而,我们发现这种严格的几何约束会引发“正交性困境”:刚性的参数隔离阻碍了语义相关任务间共享表征的迁移与积累。在本工作中,我们提出了一种新的无需回放的方法,称为整合与解耦LoRA(CoDe-LoRA),用于大语言模型的持续学习。CoDe-LoRA将学习过程分解为整合通用知识与解耦任务特定知识。为实现此目标,CoDe-LoRA利用自适应零空间投影机制和语义路由来平衡知识积累与任务特定适配。在四个骨干网络和三个持续学习基准上的实验结果表明,CoDe-LoRA取得了最佳平均准确率。我们的代码可在该https链接获取。
英文摘要
Continual learning (CL) is essential for Large Language Models (LLMs) to sequentially adapt to evolving tasks. To mitigate catastrophic forgetting, recent advances implement low-rank adaptation with orthogonal projections (e.g., O-LoRA) to isolate task parameters. However, we reveal that such strict geometric constraints trigger an "Orthogonality Dilemma": rigid parameter isolation impedes the transfer and accumulation of shared representations across semantically related tasks. In this work, we propose a new replay-free method, called Consolidation and Decoupling LoRA (CoDe-LoRA), for CL of LLMs. CoDe-LoRA disentangles the learning process into Consolidating Universal Knowledge and Decoupling Task-Specific Knowledge. To achieve this, CoDe-LoRA leverages an adaptive null space projection mechanism and semantic routing to balance knowledge accumulation with task-specific adaptation. Experimental results across four backbones and three CL benchmarks show that CoDe-LoRA achieves the best average accuracy. Our code is available at https://github.com/Estrellajer/CoDe-LoRA.
CommentsAccepted to EMNLP 2026 (Main Conference)