arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26820cs.CV

LLaVAFlow:用于参数高效多模态微调的潜在对齐流保留方法

LLaVAFlow: Preserving Latent Alignment Flow for Parameter-Efficient Multimodal Fine-Tuning

Muyao Yuan, Muyan Jiao, Jiangyong Ying, Weizhan Zhang, Yuanhong Zhang, Lan Ma, Yuan Gao, Haipeng Du

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态大语言模型微调的灾难性遗忘问题,本文提出即插即用的LLaVAFlow框架,通过信息论蒸馏保留跨模态对齐流,提升下游任务性能与泛化能力。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)虽具备强泛化能力,但针对下游任务的视觉指令微调不可避免会引发灾难性遗忘,损害整体泛化性能。现有方法虽通过调控权重更新减少遗忘,却忽视了MLLMs中基础的跨模态对齐。基于前期研究及本文观察,本文认为跨模态对齐隐含于信息压缩轨迹中。为保留该轨迹中嵌入的对齐流,本文提出LLaVAFlow——一种基于信息论的蒸馏框架:首先,压缩提取的关系与MLLM嵌入间的互信息,促使可学习模块生成利于下游任务的精细化对齐流;其次,最大化预训练与微调后MLLMs的提取对齐流间的互信息,实现紧凑对齐信息的迁移。大量实验表明,LLaVAFlow是一种即插即用的有效框架,可保留对齐流并提升下游性能与泛化能力。

英文摘要

While Multimodal Large Language Models (MLLMs) exhibit strong generalization, visual instruction tuning for downstream tasks inevitably causes catastrophic forgetting, impairing overall generalization. While existing methods regulate weight updates to reduce forgetting, they overlook the fundamental cross-modal alignment in MLLMs. Based on prior work and our observations, we argue that cross-modal alignment is implicitly captured in the information-compression trajectory. To preserve the alignment flow embedded in the trajectory, we propose LLaVAFlow, an information-theoretic distillation framework. First, we compress the mutual information between the extracted relations and MLLM embeddings, encouraging a learnable module to produce a refined alignment flow that benefits downstream tasks. Second, we maximize the mutual information between the extracted alignment flows of the pretrained and fine-tuned MLLMs, enabling the transfer of compact alignment information. Extensive experiments show that LLaVAFlow is an effective plug-and-play framework that preserves alignment flow and enhances both downstream performance and generalization.

发表机构

  • MOEKLINNS
  • China Telecom(中国电信)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑