arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15085cs.CL

为何视觉无法作为通用桥梁:修正多语言多模态大语言模型(MLLMs)中的模态异步性

Why Vision Fails as a Universal Bridge: Rectifying Modality Asynchrony in Multilingual MLLMs

Yihang Du, Juhao Liang, Zhengzhao Lai, Siyu Li, Yan Hu

首次发表
浏览论文内容

中文总结 AI 辅助

针对多语言MLLMs在非英语视觉推理中的性能下降问题,该研究发现其源于“幽灵锚点”模态异步性,提出ANCHOR训练框架,经实验验证可提升跨语言视觉推理性能。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)尽管仅文本的主干模型具备强大的多语言能力,但在非英语视觉推理任务中表现出显著的性能下降。虽然仅文本模型的机制证据表明,非英语输入会被路由到以英语为中心的潜在空间,但这一现象的多模态影响尚未得到探索。通过严格的机制分析,我们发现了“幽灵锚点(Ghost Anchor)”现象:一种时序模态异步性,即语言向英语语义流形的翻译在模型早期层完成,而视觉语义化仍未成熟,导致视觉信号在早期对齐窗口中物理存在但功能不可见。为修正这一问题,我们提出了训练框架ANCHOR,其采用主动视觉锚定(Proactive Visual Anchoring,PVA)来加速早期视觉语义的出现,确保视觉表示主动引导语言翻译。机制干预实验证实,ANCHOR成功恢复了早期翻译过程中视觉信号的因果影响。此外,在XMMMU、MaXM和CVQA上开展的大量实验表明,ANCHOR始终优于标准基线,在微调语言和零样本语言上均实现了稳健的视觉推理。

英文摘要

Multimodal large language models (MLLMs) exhibit substantial performance degradation in non-English visual reasoning, despite the strong multilingual competence of their text-only backbones. While mechanistic evidence from text-only models suggests that non-English inputs are routed through an English-centric latent space, the multimodal implications of this phenomenon remain unexplored. Through rigorous mechanistic analysis, we identify the \textbf{Ghost Anchor} phenomenon: a temporal modality asynchrony where linguistic translation to the English semantic manifold completes in early layers, while visual semanticization remains immature. Consequently, visual signals are physically present yet functionally invisible during the early alignment window. To rectify this, we propose \textbf{ANCHOR}, a training framework employing Proactive Visual Anchoring (PVA) to accelerate early visual semantic emergence, ensuring visual representations proactively guide linguistic translation. Mechanistic interventions confirm that ANCHOR successfully restores the causal influence of visual signals during early translation. Furthermore, extensive experiments on XMMMU, MaXM, and CVQA demonstrate that ANCHOR consistently outperforms standard baselines, achieving robust visual reasoning across both fine-tuned and zero-shot languages.

发表机构

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Shenzhen Loop Area Institute(深圳河套学院)
  • National Health Data Institute (Shenzhen)(国家健康数据研究院(深圳))

机构由 AI 辅助整理,请以论文原文为准。

↑