arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TokenSwap:多模态大语言模型中模态差距的基准测试与缩减

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

Andong Hua, Colton Bishop, Igor Mordatch, Arian Hosseini, Jindong Gu, Aleksandra Faust, Rebecca Roelofs, Yao Qin

arXiv 2607.28640首次发表:更新:

发表机构

University of California, Santa Barbara; Google DeepMind(加州大学圣巴巴拉分校; 谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出TokenSwap方法构建基准TokenSwap-Bench,发现42个多模态大语言模型普遍存在模态差距,推理模型差距更小,训练中引入TokenSwap可有效缓解该差距。

AI 中文摘要

多模态大语言模型(Multimodal large language models, MLLMs)在跨模态语义等价输入下应生成一致响应,但本文观察到模型预测存在系统性差异,将这种文本与多模态语义等价输入下的性能差异定义为模态差距。本文提出TokenSwap方法,通过用语义对齐的图像替换文本概念构造输入,形成视觉标记与文本标记交错的序列;基于TokenSwap将MMLU等现有文本基准转换为图像交错版本,得到TokenSwap-Bench。在42个MLLMs上的实验显示,模态差距普遍存在:从纯文本输入切换到图像交错输入时,性能下降4.2%至47.4%,模型平均下降19.6%±3.3%;推理模型的平均差距为10.1%,显著小于非推理模型的25.5%;仅靠提示策略或训练计算规模扩大无法可靠缩小该差距;最后验证训练中引入TokenSwap可有效缓解模态差距,同时保持出色的纯文本及视觉-语言性能。

英文摘要

Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal variations. Specifically, we define the modality gap as the difference in model performance under semantically equivalent textual and multimodal inputs. We introduce TokenSwap, a method that constructs such inputs by replacing textual concepts with semantically aligned images, resulting in sequences where visual tokens are interleaved with text tokens. Based on TokenSwap, we transform existing text-based benchmarks such as MMLU into image-interleaved counterparts, resulting in TokenSwap-Bench. Across 42 MLLMs, we observe a pervasive modality gap, with performance decreasing by 4.2% to 47.4% when moving from text-only to image-interleaved inputs, averaging 19.6% +/- 3.3% across models. Notably, we observe that reasoning models exhibit consistently smaller gaps, achieving an average gap of 10.1% compared to 25.5% for non-reasoning models. In contrast, neither prompting strategies nor scaling training compute alone reliably reduces the modality gap. Finally, we demonstrate that incorporating TokenSwap during training effectively mitigates this gap while preserving strong text-only and vision-language performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑