广义多模态基础模型
Fusion Anything: A Generalized Multimodal Foundation Model
- Tianjin University(天津大学)
- Beijing University of Posts and Telecommunications(北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出广义多模态基础模型,通过合成多模态数据训练,编码可迁移相关性,在18个数据集上无需适配即可媲美专门模型。
AI中文摘要:
利用多模态数据进行预测在多种场景中被广泛使用。现有的多模态融合模型一旦部署,只能处理预定义模态(如视觉、文本和音频)和单一任务,难以快速适应新的下游应用。因此,一个自然而却相当激进的问题出现了:是否存在一个通用的多模态融合模型,可以应用于任意模态组合和任意预测任务。我们认为,一个统一的多模态融合模型不应依赖于特定模态,而应编码多模态相关性的可迁移模式。为此,我们提出了一种简单有效的学习范式,基于生成具有多样因果结构的大规模合成多模态数据集进行训练,这些因果结构正式刻画了现实世界中多模态数据的生成过程。在此框架基础上,我们提出了广义多模态基础模型,这是一个用于广义多模态学习的统一基础模型。通过构建具有多样相关模式的大规模合成多模态数据集,我们的模型在训练期间编码可迁移的多模态相关性,并在推理期间通过上下文示例激活适当的关联。在涵盖12种模态和11种预测任务的18个真实世界数据集上进行的大量实验表明,我们的模型在无需任务特定适配的情况下,达到了与专门模型相当的性能。
英文摘要:
Making prediction with multimodal data is widely used in diverse scenarios. Existing multimodal fusion models, once deployed, can only handle predefined modalities (e.g., vision, text and audio) and single task, making it difficult to quickly adapt to new downstream applications. Therefore, a natural yet aggressive question arises - whether there exists a general multimodal fusion model that can be applied to arbitrary modality combinations and arbitrary prediction tasks. We argue that a unified multimodal fusion model should not depend on specific modalities and should instead encode transferable patterns of multimodal correlation. To this end, we propose a simple and effective learning paradigm based on training on large-scale synthetic multimodal datasets generated with Structural Multimodal Causal Models (SMCMs), which formally characterizes the generative processes of real-world multimodal data. Building on this framework, we propose the Fusion Anything Model (FAM), a foundation model for generalized multimodal data fusion. By constructing large-scale synthetic multimodal data with diverse correlation patterns, our model encodes transferable multimodal correlations during training and activates appropriate associations through in-context examples during inference. Extensive experiments on 18 real-world datasets spanning 12 modalities and 11 prediction tasks demonstrate that our model achieves competitive performance with specialized models without task-specific adaptation.