发表机构
National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有多模态大语言模型的图像级对齐存在指称歧义的问题,提出多模态代码切换(MMCS)范式,构建含77.3万样本的数据集,仅用5万样本即可匹配或超越60万图像-文本对训练的模型,提升了视觉基础与感知能力。
AI 中文摘要
现有多模态大语言模型(Multimodal Large Language Models, MLLMs)主要依赖图像-文本对进行模态对齐预训练,将全局图像表示映射为长文本描述。然而这种图像级对齐存在指称歧义:模型难以从全局表示中推断多个视觉对象与文本实体之间的对应关系,导致数据效率低下和语义基础不佳。为解决该问题,我们提出多模态代码切换(MultiModal Code-Switching, MMCS),一种提供显式对象级监督的新型预训练范式。受代码切换的语言现象启发,MMCS通过用对应视觉对象替换文本实体来交织视觉与语言,强化局部视觉-语言基础。我们进一步开发可扩展的数据合成流水线,生成包含77.3万个样本的预训练数据集,具备准确的对象-实体对应关系。实验表明MMCS数据效率极高:仅用5万个样本,其性能可匹配或超越在60万个图像-文本对上训练的模型;此外,MMCS在不同模型规模下均持续提升视觉基础与感知能力。
英文摘要
Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, leading to data inefficiency and suboptimal semantic grounding. To address this, we propose MultiModal Code-Switching (MMCS), a novel pretraining paradigm that provides explicit object-level supervision. Inspired by the linguistic phenomenon of code-switching, MMCS interleaves vision and language by replacing textual entities with their corresponding visual objects, enforcing local vision-language grounding. We further develop a scalable data synthesis pipeline to generate a pretraining dataset of 773K samples with accurate object-entity correspondences. Experiments show that MMCS is highly data-efficient: with only 50K samples, it matches or surpasses models trained on 600K image-text pairs. Furthermore, MMCS consistently improves visual grounding and perception capabilities across varying model scales.