用于持续指令微调的渐进式多模态对齐
Progressive Multimodal Alignment for Continual Instruction Tuning
浏览论文内容
中文总结 AI 辅助
针对多模态持续指令微调中投影器级遗忘问题,提出渐进式多模态对齐框架PMA,以亚线性参数增长平衡稳定性与可塑性,在多基准实验中提升了现有方法性能且适配多种MLLM主干。
中文摘要 AI 辅助
多模态大语言模型(Multimodal Large Language Models, MLLMs)依赖一个投影器将视觉表示与语言嵌入空间对齐,这是跨模态理解的核心。然而在多模态持续指令微调(Multimodal Continual Instruction Tuning, MCIT)中,视觉分布的变化和指令语义的演变会导致该共享投影器发生漂移,进而引发投影器级别的遗忘,这一问题在主要聚焦于大语言模型(LLM)主干的方法中被极大忽略。我们提出渐进式多模态对齐(Progressive Multimodal Alignment, PMA)框架,该框架可让投影器在保留先前学习到的对齐的同时持续适应。PMA通过轻量型表示描述符检测多模态分布偏移,并仅在需要时逐步扩展投影器专家;可扩展路由器基于多模态特征整合专家输出,同时保留原始预训练投影器作为稳定对齐锚点。该渐进式机制以亚线性参数增长平衡了稳定性与可塑性,可作为与方法无关的附加组件应用于现有MCIT方法。在两个最新MCIT基准上的大量实验表明,结合PMA缓解投影器级遗忘后,相比现有最先进方法取得了一致的性能提升;此外,PMA可在不同MLLM主干上扩展,展现出稳健且广泛适用的MCIT性能。
英文摘要
Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cross-modal understanding. In Multimodal Continual Instruction Tuning (MCIT), however, shifting visual distributions and evolving instruction semantics cause this shared projector to drift, leading to projector-level forgetting, an issue largely overlooked by methods that focus primarily on the LLM backbone. We introduce Progressive Multimodal Alignment (PMA), a framework that enables the projector to adapt continually while preserving previously learned alignment. PMA detects multimodal distribution shifts via a lightweight representation descriptor and progressively expands projector experts only when needed. An expandable router integrates expert outputs based on multimodal features, while the original pretrained projector is retained as a stable alignment anchor. This progressive mechanism balances stability and plasticity with sub-linear parameter growth and serves as a method-agnostic add-on to existing MCIT approaches. Extensive experiments on two recent MCIT benchmarks demonstrate that mitigating projector-level forgetting yields consistent gains over prior state-of-the-art methods when combined with PMA. Moreover, PMA scales across diverse MLLM backbones, demonstrating robust and broadly applicable MCIT performance.
发表机构
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
- Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences(中国科学院脑科学与智能技术卓越创新中心)
- Kyoto University(京都大学)
- Migu Culture Technology Co.,Ltd.(咪咕文化科技有限公司)
- State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与类脑智能技术国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。