发表机构
Ramen VR; University of California, Berkeley(拉面VR公司; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究多模态大语言模型适配新模态是否需微调主干,经实验发现仅训练投影器即可实现强多模态性能,还能避免联合训练导致的语言模型能力漂移,且训练样本吞吐量约为联合训练的两倍,通过多类基准验证了结论。
AI 中文摘要
多模态大语言模型(MLLM)的典型训练流程包括同时适配语言模型主干,以及主干与特定模态编码器之间的投影器。本文探究将MLLM的主干微调以适配新模态是否必要。通过在3D MLLM上开展实验,研究发现仅训练投影器即可实现与现有基准模型、以及采用相同编码器和主干的联合训练MLLM相当的强多模态性能。研究还表明,联合训练会导致语言模型现有能力出现不理想的漂移,而仅训练投影器从定义上可避免该问题。此外,仅训练投影器的训练样本吞吐量约为联合训练的两倍。研究通过3D分类、字幕生成基准,以及评估语言、视觉和空间推理能力的标准基准,在不同语言模型主干上验证了上述发现。
英文摘要
The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder. We ask whether fine-tuning the backbone of an MLLM is necessary to adapt it to a new modality. Through experiments on 3D MLLMs, we find that training only the projector is sufficient to achieve strong multimodal performance relative to existing baseline models and our jointly trained MLLMs with the same encoder and backbone. We also show that joint training leads to undesirable drift in existing capabilities of the language model, which projector-only training avoids by definition. Furthermore, projector-only training has approximately twice the training sample throughput of joint training. We validate our findings across different language model backbones via 3D classification and captioning benchmarks as well as standard benchmarks evaluating language, vision, and spatial reasoning capabilities.