arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态大语言模型中依赖架构的融合路径

Architecture-Dependent Fusion Pathways in MLLMs

Hebao Zhu, Dongxia Wu

arXiv 2610.03289首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过分析拼接与原生多模态架构的MLLMs,揭示两种不同的融合路径,为多模态融合提供机制性视角及架构感知诊断方法。

AI 中文摘要

多模态大语言模型(MLLMs)在视觉-语言任务中表现出强劲性能,然而视觉与文本信息在各层之间融合的内部机制仍未被充分理解。我们研究了来自两种架构范式的代表性MLLMs:拼接架构和原生多模态架构。我们进行了三项逐步关联的分析:对齐解耦识别哪些模态发生变化,注意力路由与熵刻画跨模态信息如何分布,以及内在维度考察融合如何重塑特征空间。此外,我们进行因果干预实验以验证所得解释。作为补充分析,我们使用视觉CKA检验柏拉图式表征假说。综合这些分析,我们揭示了两种不同的融合路径:拼接模型遵循文本优先、视觉滞后的路径,而原生模型表现出更早的视觉-文本协同适应及特征空间重组。这项工作为理解多模态融合提供了机制性视角,并支持对多模态表征进行架构感知的诊断。

英文摘要

Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood. We investigate representative MLLMs from two architectural paradigms: concatenation architectures and native multimodal architectures. We conduct three progressively connected analyses: alignment decoupling identifies which modality changes, attention routing and entropy characterize how cross-modal information is distributed, and intrinsic dimensionality examines how fusion reshapes feature spaces. Separately, we perform causal intervention experiments as a validation of the resulting interpretation. As a supplementary analysis, we use visual CKA to examine the Platonic Representation Hypothesis. Together, these analyses reveal two distinct fusion pathways: concatenation models follow a text-first, vision-later pathway, whereas native models exhibit earlier visual-textual co-adaptation and feature-space reorganization. This work provides a mechanistic perspective for understanding multimodal fusion and supports architecture-aware diagnostics of multimodal representations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑