用于多模态理解与生成的解耦式视觉-语言系统
Decoupled Vision-Language System for Multimodal Understanding and Generation
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- Peng Cheng Laboratory(鹏城实验室)
- Li Auto(理想汽车)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出多模态大语言模型Libra的解耦式视觉-语言架构,通过开关注意力与FFN模块实现自模态与跨模态解耦,在理解与生成任务上均取得优异性能。
AI中文摘要:
我们提出了一种用于多模态大语言模型(MLLM)的新架构设计Libra,该模型兼具多模态理解与生成能力。Libra架构包含一个视觉系统和一个语言系统,通过跨模态桥接连接。此设计将自模态建模与跨模态解耦,使每种模态能够学习其独特表示,同时保持有效的跨模态理解。解耦主要通过开关注意力模块和开关前馈网络(FFN)模块实现,这两个模块可动态为自模态建模和跨模态交互场景路由计算流。我们在两个重要场景中评估其有效性:仅用于理解的图像到文本场景的Libra-1,以及统一图像到文本理解与文本到图像生成的Libra-2。除架构设计外,我们还讨论了分词、位置编码和监督方面的各种改进。实验表明,专用的Libra设计可实现多模态理解与生成的相互提升,在理解和生成基准上均取得了优异性能。
英文摘要:
We introduce a new architecture design for multimodal large language models (MLLMs), Libra, capable of both multimodal understanding and generation. Libra architecture contains one vision system and one language system, connected by cross-modal bridges. This design decouples self-modal modeling and cross-modal interaction, enabling each modality to learn its unique representations while maintaining effective cross-modal comprehension. The decoupling is mainly achieved in a switch attention module and a switch FFN module, which dynamically routes the computation flow for self-modal modeling and cross-modal interaction scenarios. We evaluate the effectiveness in two important settings: \textbf{Libra-1} for the understanding-only image-to-text setting, and \textbf{Libra-2} for unified image-to-text understanding and text-to-image generation. In addition to the architecture design, we discuss various improvements on tokenization, positional encoding, and supervision. Experiments demonstrate that the dedicated Libra design enables mutual improvements on multimodal understanding and generation, achieving strong performance on both understanding and generation benchmarks.