发表机构
Gaoling School of Artificial Intelligence, Renmin University of China; Beijing Key Laboratory of Research on Large Models and Intelligent Governance; Beihang University; Beijing University of Posts and Telecommunications; Shanghai Artificial Intelligence Laboratory; AresoX(中国人民大学高瓴人工智能学院; 北京大模型与智能治理重点实验室; 北京航空航天大学; 北京邮电大学; 上海人工智能实验室; AresoX)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Gestalt,一种基于多模态交互金字塔的大型多模态模型,采用统一离散扩散框架和交互分区架构,通过可学习交互令牌实现跨模态整合,在图像生成、多模态理解和文本评估中表现优异。
AI 中文摘要
在本文中,我们提出了Gestalt,一种围绕多模态交互构建的大型多模态模型的新范式。尽管取得了快速进展,大型多模态模型正面临瓶颈:现有方法主要侧重于容纳更多模态,却忽视了每种模态的独特特征以及它们之间的关系。受人类多感官感知的多阶段特性的启发,我们提出了一个多模态交互金字塔,将多模态建模组织为从模态特定处理、跨模态对齐到更深层次多模态整合的递进过程。在该金字塔的指导下,Gestalt采用统一的离散扩散框架和交互分区架构,通过可学习的交互令牌来调节跨模态交换与整合。该金字塔还构建了其数据组织和训练策略。在图像生成、多模态理解和纯文本评估方面的强劲表现表明,Gestalt在保留模态特定信息的同时显著改善了跨模态整合,有效利用了基于扩散的多模态模型的优势,为迈向统一多模态智能提供了一条有前景的路径。
英文摘要
In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal models are reaching a bottleneck: existing approaches focus primarily on accommodating additional modalities while overlooking the distinct characteristics of each modality and the relations among them. Motivated by the multistage property of human multisensory perception, we propose a multimodal interplay pyramid that organizes multimodal modeling as a progression from modality-specific processing, through cross-modal alignment, to deeper multimodal integration. Guided by this pyramid, Gestalt adopts a unified discrete diffusion framework and an interplay-partitioned architecture, with learnable interplay tokens mediating cross-modal exchange and integration. The pyramid also structures its data organization and training strategy. Strong performance across image generation, multimodal understanding, and text-only evaluation shows that Gestalt significantly improves cross-modal integration while preserving modality-specific information, effectively harnessing the strengths of diffusion-based multimodal models and offering a promising path toward unified multimodal intelligence. Project page: https://GeWu-Lab.github.io/Gestalt.
Comments17 pages, 7 figures