C3-UniMM:通过超级对齐与共享解码空间实现因果循环一致性的统一多模态建模
C3-UniMM: Causal Cycle-Consistent Unified Multimodal Modeling via Super Alignment and Shared Decoding Space
- Hunan University(湖南大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出C3-UniMM框架,通过结构化隐式因果图(SLCG)与统一解码空间解决现有统一多模态模型的跨模态结构一致性问题,在多任务上性能优于现有基线。
AI中文摘要:
统一多模态模型旨在实现任意模态间的任意到任意的理解与生成,但现有方法主要依赖对隐式统计关联的建模,缺乏跨模态结构一致性约束,这一缺陷引发了语义漂移、组合泛化能力差、干预下不稳定等严重问题。本文提出C3-UniMM,一种基于因果循环一致性与超级对齐的统一多模态建模框架:引入结构化隐式因果图(SLCG)作为共享跨模态语义空间,设计统一多模态编码块,使理解与生成能在同一因果语义结构中协同优化;还提出统一解码空间,以在跨模态生成过程中强化结构保留与语义可逆性。理论分析表明,该方法显著提升了跨模态映射的可逆性与机制不变性;在多个理解、生成及组合泛化任务上的大量实验结果显示,C3-UniMM的性能大幅优于现有统一多模态基线模型。
英文摘要:
Unified Multimodal Models aim to achieve any-to-any understanding and generation across arbitrary modalities. However, existing methods primarily rely on modeling implicit statistical correlations and lack cross-modal structural consistency constraints. This deficiency leads to profound issues, including semantic drift, poor compositional generalization, and instability under interventions. In this paper, we propose C3-UniMM, a unified multimodal modeling framework based on Causal Cycle Consistency and Super Alignment. Specifically, we introduce a Structured Latent Causal Graph (SLCG) as a shared cross-modal semantic space and design unified multimodal encoding blocks, enabling understanding and generation to be synergistically optimized within the identical causal semantic structure. Furthermore, we propose a Unified Decoding Space to enforce structural preservation and semantic invertibility during the cross-modal generation process. Theoretical analyses demonstrate that our approach significantly enhances both the invertibility and mechanism invariance of cross-modal mappings. Extensive experimental results across multiple understanding, generation, and compositional generalization tasks indicate that C3-UniMM substantially outperforms existing unified multimodal baselines.