发表机构
Southern University of Science and Technology; Microsoft Research Asia; Shanghai University of Finance and Economics; Peking University(南方科技大学; 微软亚洲研究院; 上海财经大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对多模态大语言模型的英语中心化问题,提出VFA框架,通过组合任务向量实现多语言增强与视觉对齐解耦,仅用少量数据就提升了多语言模型性能并保留原有能力。
AI 中文摘要
多模态大语言模型发展迅速,但多数仍以英语为中心,因为高质量非英语文本-图像监督数据稀缺且成本高昂,限制了多语言多模态指令调优的扩展。尽管多语言文本数据丰富,单纯的文本微调会破坏视觉-语言对齐并导致灾难性遗忘。我们提出无视觉适配(Vision-Free Adaptation,VFA)框架,该框架通过在共享大语言模型(LLM)主干上组合互补任务向量,将多语言语言增强与视觉解耦。具体而言,我们在多语言文本数据上微调基础LLM以得到多语言任务向量,再将其与某多模态大语言模型(MLLM)的视觉对齐任务向量合并。在6个多语言多模态基准上对5种MLLM开展的实验显示,VFA在保留通用多模态和纯文本能力的同时实现了一致提升;此外,仅使用不到2%的文本数据,VFA就缩小了与完全多模态训练模型的差距,展现出数据效率。
英文摘要
Multimodal large language models have advanced rapidly, yet most remain English-centric, as scaling multilingual multimodal instruction tuning is limited by the scarcity and high cost of high-quality non-English image-text supervision. Although multilingual text data is abundant, naive textual fine-tuning can disrupt vision-language alignment and induce catastrophic forgetting. We propose Vision-Free Adaptation (VFA), a framework that decouples multilingual language enhancement from visual alignment by composing complementary task vectors over a shared LLM backbone. Specifically, we fine-tune a base LLM on multilingual text data to derive a multilingual task vector, which is then merged with the vision-aligned task vector of an MLLM. Experiments on five MLLMs across six multilingual multimodal benchmarks show consistent improvements while preserving both general multimodal and text-only capabilities. Moreover, using less than 2% of the text data, VFA narrows the gap to the fully multimodal-trained model, demonstrating its data efficiency.