发表机构
Seoul National University(首尔国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多模态持续指令微调中存在的问题,提出SiGMA框架,通过训练时符号引导自适应调优、推理时符号引导合并减轻负面干扰,在相关基准实验中显著减少干扰且性能优于现有方法。
AI 中文摘要
多模态持续指令微调(MCIT)对于使多模态大语言模型(MLLMs)适应不断演变的下游任务序列至关重要。先前方法大多采用专家混合或扩展合并方法,主要关注灾难性遗忘,但在推理过程中仍存在负面干扰,新学习的更新会覆盖有用的先验知识并降低整体性能。为解决此问题,我们提出了SiGMA(符号引导合并与适配),这是一个简单而有效的框架,通过训练期间的符号引导自适应调优和推理时的符号引导合并两个组件来减轻负面干扰。符号引导自适应调优减少与过去知识的冲突,并以最小的漂移学习当前任务,减轻严重遗忘。符号引导合并通过有选择地缩放显著参数进一步改进整合,以保留和放大有用的特定任务知识。在UCIT和DCL基准上的实验表明,SiGMA显著减少负面干扰并优于现有MCIT方法。我们的代码可在SiGMA获取。
英文摘要
Multimodal Continual Instruction Tuning (MCIT) is crucial for adapting Multimodal Large Language Models (MLLMs) to evolving a sequence of downstream tasks. Prior methods mostly utilize Mixture of Experts or expansion merge approach, primarily focusing on catastrophic forgetting, yet they still suffer from negative interference during inference, where newly learned updates overwrite useful prior knowledge and degrade overall performance. To address this, we propose SiGMA (Sign Guided Merging and Adaptation), a simple yet effective framework that mitigates negative interference with two components: sign guided adaptive tuning during training and sign guided merging at inference. Sign guided adaptive tuning reduces collisions with past knowledge and learns the current task with minimal drift, mitigating severe forgetting. Sign guided merging further improves consolidation by selectively scaling salient parameters to preserve and amplify useful task specific knowledge. Experiments on UCIT and DCL benchmarks show that SiGMA significantly reduces negative interference and outperforms state of the art MCIT methods. Our code is available at SiGMA.
CommentsAccepted at ECCV 2026