发表机构
School of Software Technology, Zhejiang University; College of Intelligent Robotics and Advanced Manufacturing, Fudan University; School of Artificial Intelligence, Sun Yat-Sen University; Graduate School of Information Science, Hokkaido University(浙江大学软件学院; 复旦大学智能机器人与先进制造学院; 中山大学人工智能学院; 北海道大学信息科学研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大语言模型微调时的灾难性遗忘问题,提出无需修改架构的AWARe方法,通过激活加权控制参数更新,在保留上游能力的同时提升下游性能。
AI 中文摘要
多模态大语言模型(MLLMs)因大规模多模态预训练展现出强大的泛化与推理能力,但对其进行下游任务微调常导致灾难性遗忘——新学习的任务特定知识会降低先前获取的能力,该问题源于新任务的梯度更新会覆盖对先前知识至关重要的参数,限制了MLLMs的实际部署。为解决此挑战,本文提出激活加权自适应保留(AWARe),这是一种通过基于激活模式动态控制参数更新来缓解灾难性遗忘的微调方法。AWARe为参数分配基于激活的重要性分数,选择性冻结对保留先前能力至关重要的参数,同时允许重要性较低的参数适配新任务;重要的是,AWARe无需修改模型架构,确保与现有推理引擎兼容。大量实验表明,与现有方法相比,AWARe能有效保留上游能力,同时实现更优的下游性能,代码可在指定URL获取。
英文摘要
Multimodal Large Language Models (MLLMs) exhibit strong generalization and reasoning abilities due to large-scale multimodal pre-training. However, fine-tuning these models on downstream tasks often leads to catastrophic forgetting, where newly learned task-specific knowledge degrades previously acquired capabilities. This issue arises because gradient updates for new tasks overwrite parameters critical to prior knowledge, limiting the practical deployment of MLLMs. To address this challenge, we propose Activation-Weighted Adaptive REtention (AWARe), a fine-tuning method that mitigates catastrophic forgetting by dynamically controlling parameter updates based on activation patterns. AWARe assigns activation-based importance scores to parameters, selectively freezing those essential for preserving prior capabilities while allowing less important parameters to adapt to new tasks. Importantly, AWARe operates without modifying model architectures, ensuring compatibility with existing inference engines. Extensive experiments demonstrate that AWARe effectively preserves upstream capabilities while achieving superior downstream performance compared to existing methods. Code is available at https://github.com/kaln27/AWARe.