AI 中文总结
研究任意到任意多模态建模,提出仅解码器的对称多模态建模方法Modus,无需特定模态组件,支持多种应用,在各基准测试中用单模型展现强大性能且与基线竞争,材料开源。
AI 中文摘要
任意到任意模型可在单个网络中根据其他模态的任意组合预测任何模态,此方法用于多模态视觉和视觉语言模型,在生态学和天文学等科学领域也越来越常用。现有任意到任意模型通常使用编码器-解码器或扩散架构从头开始训练,影响其性能且无法利用强大的仅解码器预训练模型。本文研究仅解码器的任意到任意多模态建模,该方法对称对待所有模态,支持任意模态作为输入和输出,无需特定模态头、损失或任务管道。由此产生的名为Modus的模型可支持一系列应用,如通过中间模态进行链式生成或通过用另一个生成的模态对模型自身输出进行评分来进行跨模态自验证。Modus展现出强大的开箱即用性能,在各种基准测试中使用单个模型与专业和多任务基线竞争。所有材料在该https网址开源。
英文摘要
Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any-to-any multimodal modeling, which treats all modalities symmetrically and supports arbitrary modalities as inputs and outputs without modality-specific heads, losses, or task pipelines. Because every modality is both an input and an output of the same model, the resulting model, named Modus, can support a range of applications, such as chained generation through intermediate modalities or cross-modal self-verification by scoring the model's own outputs with another generated modality. Modus demonstrates strong out-of-the-box performance and is competitive with specialist and multitask baselines using a single model across various benchmarks. All materials are open-sourced at https://modus-multimodal.epfl.ch/.
CommentsAccepted at ICML 2026. Project page: https://modus-multimodal.epfl.ch