arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2311.17963cs.CV

M$^{2}$Chat:赋能VLM实现多模态LLM交错文本-图像生成

M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation

  • The Hong Kong University of Science and Technology(香港科学与技术大学)
  • Waseda University(早稻田大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiaowei Chi, Junbo Qi, Rongyu Zhang, Shanghang Zhang, Qifeng Liu, Yike Guo

更新

AI总结:

提出M$^{2}$Chat统一多模态LLM框架,通过M$^{3}$Adapter融合视觉与语义特征并采用门控策略及两阶段微调,实现高质量交错文本-图像生成,超越现有基准。

AI中文摘要:

尽管当前的LLM聊天机器人(如GPT-4V)弥合了人类指令与视觉表示之间的差距,实现了文本-图像生成,但它们仍然缺乏高效的对齐方法,以在多个下游任务上实现高保真性能。在本文中,我们提出了M$^{2}$Chat,一个新颖的统一多模态LLM框架,用于在各种场景下生成交错的文本-图像对话。具体来说,我们提出了M$^{3}$Adapter,它高效地整合了来自多模态提示的细粒度低级视觉信息和高级语义特征。在良好对齐的融合特征之上,M$^{3}$Adapter定制了一种可学习的门控策略,以自适应地平衡模型在不同任务上的创造性和一致性。此外,为了进一步增强M$^{3}$Adapter的有效性,同时保持语义上下文理解的连贯性,我们引入了一种两阶段的M$^{3}$FT微调策略。该策略分别优化用于图像-文本对齐和视觉指令的不相交参数组。大量实验表明,我们的M$^{2}$Chat在各种基准测试中超越了最先进的同类模型,展示了其在交错生成、故事讲述和多模态对话系统中的卓越能力。演示和代码可在https://mattie-e.github.io/M2Chat.github.io获取。

英文摘要:

While current LLM chatbots like GPT-4V bridge the gap between human instructions and visual representations to enable text-image generations, they still lack efficient alignment methods for high-fidelity performance on multiple downstream tasks. In this paper, we propose \textbf{$M^{2}Chat$}, a novel unified multimodal LLM framework for generating interleaved text-image conversation across various scenarios. Specifically, we propose an $M^{3}Adapter$ that efficiently integrates granular low-level visual information and high-level semantic features from multi-modality prompts. Upon the well-aligned fused feature, $M^{3}Adapter$ tailors a learnable gating strategy to balance the model creativity and consistency across various tasks adaptively. Moreover, to further enhance the effectiveness of $M^{3}Adapter$ while preserving the coherence of semantic context comprehension, we introduce a two-stage $M^{3}FT$ fine-tuning strategy. This strategy optimizes disjoint groups of parameters for image-text alignment and visual-instruction respectively. Extensive experiments demonstrate our $M^{2}Chat$ surpasses state-of-the-art counterparts across diverse benchmarks, showcasing its prowess in interleaving generation, storytelling, and multimodal dialogue systems. The demo and code are available at \red{https://mattie-e.github.io/M2Chat.github.io}.

↑