arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04580cs.AIcs.CL

强助弱:多模态大语言模型中的定向跨模态对齐迁移

Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs

Hoigi Seo, Byung Hyun Lee, Minjun Kim, Dohyun Mah, Jongho Lee, Se Young Chun

首次发表
浏览论文内容

中文总结 AI 辅助

本文发现将强模态MLLM合并到弱模态MLLM可提升后者性能,源于模态与文本标记对齐增强,据此提出DCAT框架,无需微调即可实现跨模态对齐迁移,优于现有合并方法。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)通过将大语言模型(LLM)与目标模态(如视觉、视频或音频)的编码器配对,实现了强大的模态理解能力。然而,提升MLLM在特定模态上的能力通常需要在大型模态特定数据集上进行额外训练,这会产生大量的数据收集和计算成本。模型合并提供了一种替代方案,但对于数据稀缺、单样本规模大或领域特定的模态(例如音频和视频),由于同模态模型变体很少可用,模型合并往往不可行。在这项工作中,我们刻画了一个有趣的非对称现象:将一个对齐良好、数据丰富的源模态MLLM合并到数据稀缺的目标模态MLLM中,能显著提升目标模态在其自身基准上的表现。我们的理论和实证分析表明,这种提升源于更强的供体模态所诱导的模态特定标记与文本标记之间对齐的增强。具体来说,我们推导了一个互信息下界,该下界与对齐相关量单调相关,并与下游MLLM性能强相关。基于这一原理,我们提出了定向跨模态对齐迁移(DCAT),这是一个新颖的框架,将文本对齐从强大的、对齐良好的源(供体)模态迁移到弱目标(受体)模态,无需进一步微调即可提升目标模态性能。我们进一步表明,该对齐增强目标函数允许一个闭式权重空间解,该解仅需从一个小型校准集计算得出。DCAT优于现有的模型合并方法,为跨模态对齐迁移提供了一条高效路径。项目页面及代码可在 \u007b此 https URL\u007d 获取。

英文摘要

Multi-modal large language models (MLLMs) achieve strong modality understanding by pairing a large language model (LLM) with an encoder for a target modality such as vision, video, or audio. However, improving an MLLM's capability for a given modality typically requires additional training on large modality-specific datasets, incurring substantial data collection and compute costs. Model merging offers an alternative, but it is often infeasible for data-scarce, large per-sample size, or domain-specific modalities (\textit{e.g.}, audio and video), where same-modality model variants are rarely available. In this work, we characterize an intriguing asymmetric phenomenon: merging a well-aligned, data-rich source-modality MLLM into a data-scarce target-modality MLLM substantially improves the target on its own benchmarks. Our theoretical and empirical analyses show that this gain stems from enhanced alignment between modality-specific and textual tokens, induced by the stronger donor modality. Specifically, we derive a mutual-information lower bound that is monotonic in alignment-related quantities and strongly correlated with downstream MLLM performance. Building on this principle, we propose Directional Cross-modal Alignment Transfer (DCAT), a novel framework that transfers textual alignment from a strong, well-aligned source (donor) modality to a weak target (recipient) modality, boosting target-modality performance without further fine-tuning. We further show that the alignment-enhancing objective admits a closed-form weight-space solution computed from only a small calibration set. DCAT outperforms existing model-merging methods, offering an efficient path toward cross-modal alignment transfer. Project page with code is available at \url{https://seohoiki3215.github.io/DCAT_project_page}

发表机构

  • Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑