arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30210cs.CVcs.LG

多模态大语言模型中的对齐错觉

The Alignment Illusion in Multimodal Large Language Models

Hong-Han Wang, Yuntao Wang, Hu Ding

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示多模态大语言模型中视觉-文本对齐分数可能仅反映权重诱导的伪相似性,并提出主角度间隙(PA gap)作为更可靠的几何诊断指标,以准确反映视觉流与任务表现的关系。

中文摘要 AI 辅助

在多模态大语言模型(MLLMs)中,逐层的视觉-文本相似性被广泛解读为语言模型逐步将视觉内容整合到共享表示空间中的证据。这种解读基于一个假设,即标量对齐分数反映了内容层面的跨模态交互。为了检验这一假设,我们对视觉流施加了受控干预。在跨越五个家族、参数量从0.5B到72B的13个MLLM中,将投影器输出的视觉令牌替换为高斯噪声会急剧降低任务准确率,然而四种标准标量度量(CKA、SVCCA、MIR以及主角度余弦的前导值)却无法一致地将受损流与原始流区分开来。我们将这种失败称为对齐错觉,并将其追溯至共享的语言模型通路:各向异性的MLP下投影将视觉和文本令牌拉向共同的输出方向,从而产生权重诱导的对齐。由于该成分本质上是一维的,我们引入了主角度间隙(PA gap),定义为前两个主角度余弦之差,它能够将权重诱导的相似性与多方向视觉结构区分开来。在分级视觉损坏下,PA gap对任务准确率的追踪比我们所考虑的标量分数更为一致;在结构化但无关的图像下,它进一步揭示了内部几何与任务准确率分离的情形。因此,MLLM中的内部视觉-文本对齐最好被解读为语言模型内部视觉流的几何诊断指标,而非内容层面跨模态交互的直接代理,并且当通过受控任务证据进行校准时最具信息量。

英文摘要

Layer-wise visual-text similarity in Multimodal Large Language Models (MLLMs) is widely interpreted as evidence that the language model progressively integrates visual content into a shared representation space. This reading rests on the assumption that scalar alignment scores reflect content-level cross-modal interaction. To test this assumption, we apply controlled interventions to the visual stream. Across 13 MLLMs from five families spanning 0.5B to 72B parameters, replacing projector-output visual tokens with Gaussian noise sharply reduces task accuracy, yet four standard scalar measures (CKA, SVCCA, MIR, and the leading principal-angle cosine) fail to consistently separate the corrupted stream from the original. We call this failure the alignment illusion and trace it to the shared language-model pathway: anisotropic MLP down-projections pull visual and text tokens toward common output directions, producing weight-induced alignment. Because this component is essentially one-dimensional, we introduce the principal-angle gap (PA gap), defined as the difference between the top two principal-angle cosines, which separates weight-induced similarity from multi-directional visual structure. Under graded visual corruption, the PA gap tracks task accuracy more consistently than the scalar scores we consider; under a structured but irrelevant image, it further exposes regimes in which internal geometry and task accuracy come apart. Internal visual-text alignment in MLLMs is therefore best read as a geometric diagnostic of the visual stream inside the language model rather than a direct proxy for content-level cross-modal interaction, and is most informative when calibrated by controlled task evidence.

发表机构

  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑