发表机构
School of Artificial Intelligence, Beihang University; Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing; Institute of Computing Technology, Chinese Academy of Sciences(北京航空航天大学人工智能学院; 未来区块链与隐私计算北京高精尖创新中心; 中国科学院计算技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLM持续学习的灾难性遗忘问题,提出SVC方法,通过奇异向量通道的选择性更新实现稳定性-可塑性平衡,在四个LLM家族及八个下游任务上表现优于现有PEFT基线。
AI 中文摘要
大语言模型(LLM)的领域特定持续适配存在灾难性遗忘风险,在获取新能力与保留预训练所学能力之间形成了根本张力。参数高效微调(PEFT)通过限制可训练参数数量缓解该问题,但现有方法缺乏决定应分配可塑性、应保留稳定性的原则性单元。我们将奇异向量通道识别为管理该权衡的自然单元,每个通道代表一种输入-输出变换,可更新以获取新知识或固定以保留预训练能力。基于此视角,我们提出SVC,一种选择性更新奇异向量通道的参数高效持续学习方法。微调前,SVC使用领域特定数据估计每个通道的适配收益,仅使用固定的公开通用领域语料作为历史激活代理以估计遗忘成本;随后通过基于膝点的成本筛选、帕累托前沿过滤及大津阈值法,基于这些分数自适应选择可训练通道。在四个LLM家族及八个下游任务上的实验结果显示,与现有PEFT基线相比,SVC能更好地保留预训练能力,同时实现强劲的下游性能。对通道评分与选择的进一步分析表明,奇异向量通道层面的选择性可塑性可实现有效的LLM持续适配。
英文摘要
Domain-specific continual adaptation of LLMs risks catastrophic forgetting, creating a fundamental tension between acquiring new capabilities and preserving those learned during pretraining. PEFT mitigates this problem by restricting the number of trainable parameters, but existing methods lack a principled unit for deciding where plasticity should be allocated and stability should be preserved. We identify the singular-vector channel as a natural unit for managing this trade-off. Each channel represents an input-output transformation, which can be updated to acquire new knowledge or fixed to preserve pretrained capabilities. Based on this perspective, we introduce SVC, a parameter-efficient continual-learning method that selectively updates Singular-Vector Channels. Before fine-tuning, SVC uses domain-specific data to estimate each channel's adaptation benefit and a fixed public general-domain corpus only as a history activation proxy for estimating forgetting cost. It then adaptively selects trainable channels based on these scores via knee-based cost screening, Pareto-front filtering, and Otsu thresholding. Experimental results across four LLM families and eight downstream tasks show that SVC better preserves pretrained capabilities while achieving strong downstream performance relative to existing PEFT baselines. Further analysis of channel scoring and selection demonstrates that selective plasticity at the singular-vector-channel level enables effective continual LLM adaptation.