arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05000cs.CVcs.LGcs.MM

迈向多模态预训练的物理学:知识流、模态协同、早期统一与方法指南

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

  • FAIR, Meta(Meta FAIR研究院)
  • Reality Labs, Meta(Meta Reality Labs)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos, Mike Lewis

AI总结:

该研究探索多模态预训练的机制,得出知识流、模态协同等四个关键见解,推导高效预训练方法,为多模态预训练的理解与扩展提供基础。

AI中文摘要:

视觉为推进基础模型提供了关键轴,推动了向原生统一多模态预训练的转变。尽管有这一势头,统一训练期间模态如何交互的设计空间和基本机制仍未得到充分探索。我们通过对多模态预训练的系统探索提供了实证清晰性。我们在合成数据集和大规模真实世界数据集上的受控实验得出了关于多模态预训练物理学的四个关键见解:(i)知识流:我们分解了语言、视觉理解和视觉生成如何跨模态传递知识,揭示了不同的影响模式和不对称性;(ii)协同与竞争:我们表明数据“复杂性”在很大程度上决定了模态是否协同,确定了促进协同的架构选择,例如共享注意力和带有特定模态前馈层的归一化,并且发现这些行为在不同视觉分词器设计中具有通用性;(iii)早期统一:从非常早期的阶段就统一模态并联合训练,被证明比后期对齐或顺序训练更有效。这一过程揭示了视觉惰性现象,即延迟整合导致模型依赖语言先验;(iv)方法指南:我们推导了高效的预训练方法指南,仅使用5%的计算预算就能实现强大的生成性能。这些核心发现随后通过在2万亿token上训练多个135亿参数的MoE模型进行了大规模验证。我们希望这项研究为理解和扩展多模态预训练提供了有原则的基础。

英文摘要:

Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data "complexity" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.

补充信息

↑