arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从零开始扩展原生多模态预训练

Scaling Native Multimodal Pre-Training From Scratch

Haoyuan Wu, Aoqi Wu, Hai Wang, Jiajia Wu, Jinxiang Ou, Bei Yu

arXiv 2607.22043首次发表:更新:

发表机构

The Chinese University of Hong Kong; LLM Department, Tencent(香港中文大学; 腾讯大语言模型部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在固定计算预算下训练基于Transformer的视觉语言模型的最优模型大小和token数量,发现语言和多模态目标扩展行为不同,推导效率前沿,还表明原生多模态预训练能促进跨模态转移,为扩展多模态基础模型奠定基础。

AI 中文摘要

虽然大语言模型展现出卓越推理能力,但仅依赖文本预训练限制了对多模态物理世界的感知。原生多模态预训练通过对多模态输入从头训练模型避免此局限,实现深度跨模态整合并减轻传统后期融合架构固有优化不对称性。然而该范式的扩展属性仍未系统表征。本文在固定计算预算下研究训练基于Transformer的视觉语言模型的最优模型大小和token数量,发现最小目标损失遵循可预测计算法则,语言和多模态目标呈现不同扩展行为,还推导了效率前沿。下游评估表明原生多模态预训练能促进跨模态转移,增强纯文本空间推理并实现强大的多模态上下文学习。总之,该实证研究为可预测地扩展多模态基础模型奠定了基础。

英文摘要

Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this limitation by training models from scratch on multimodal inputs, thereby achieving deep cross-modal integration and mitigating optimization asymmetries inherent to traditional late-fusion architectures. Despite these advantages, the scaling properties of this paradigm remain incompletely characterized. To address this gap, we investigate the optimal model size and token count for training a Transformer-based vision-language model under a fixed computational budget. Our study demonstrates that minimal objective loss adheres to a predictable compute law, whereas compute-optimal model sizes and token counts scale as power laws. Notably, language and multimodal objectives manifest distinct allocation trends. The language allocation exponents lie in a similar range across the different data mixtures. The multimodal model-allocation exponent decreases modestly with the multimodal token ratio, indicating a relative shift toward token allocation. Additionally, our scaling analysis yields a budget-compensation rule. Specifically, an additional multimodal-token budget can offset the text-efficiency penalty caused by incorporating visual information into native multimodal pre-training under a fixed compute budget. Downstream evaluations further reveal that native multimodal pre-training is associated with improved spatial reasoning and multimodal few-shot learning. Generally, this empirical research establishes the essential groundwork for predictably scaling multimodal foundation models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑