发表机构
School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics; Chengdu Everimaging Science and Technology Co., Ltd; X-Humanoid; Hong Kong Institute of AI for Science, City University of Hong Kong; Zhejiang Sci-Tech University; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; University of Electronic Science and Technology of China; Artificial Intelligence and Digital Finance Key Laboratory of Sichuan Province(西南财经大学计算机与人工智能学院; 成都依米光电科技有限公司; X-人形机器人(X-Humanoid); 香港城市大学香港人工智能科学研究院; 浙江理工大学; 中国科学院深圳先进技术研究院; 电子科技大学; 四川省人工智能与数字金融重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出HRIL方法,通过构建跨矩张量并结合Tucker分解与协同感知正则化,在多模态表示中保留协同信息,实验显示其在多模态任务上优于现有对比方法。
AI 中文摘要
自监督多模态表示学习已在多个领域取得显著成功,但由于跨模态交互的复杂性,捕获协同信息仍然具有挑战性。与单个模态间的共享信息不同,协同信息是指仅从多个模态的联合配置中产生的任务相关信号,无法从任何单独的模态中恢复。本研究聚焦于如何在多模态表示中保留此类协同信号的信息容量。关键观察是,协同信息体现在模态间的高阶统计依赖中,这为显式建模联合交互提供了原则性目标。基于这一见解,我们提出了高阶表示与信息学习(HRIL),该方法在模态嵌入上构建经验跨矩张量以表示多向交互。HRIL采用Tucker分解获取核心张量,并辅以协同感知正则化项,以防止能量集中并保留高阶耦合容量,从而捕获协同信息。在受控协同任务和真实世界基准上的实验表明,与现有多模态对比方法相比,HRIL取得了一致的改进,在以协同交互为主的任务上增益尤为显著。代码已发布在this https URL。
英文摘要
Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unlike the shared information across individual modalities, synergy arises when task-relevant signals emerge only from the joint configuration of multiple modalities and cannot be recovered from any modality in isolation. This work focuses on how to preserve the information capacity for such synergistic signals in multimodal representations. The key observation is that synergistic information is reflected in higher-order statistical dependence among modalities, which provides a principled target for explicitly modeling joint interactions. Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions. HRIL employs Tucker decomposition to obtain a core tensor, complemented by a synergy-aware regularizer that prevents energy concentration and preserves higher-order coupling capacity for synergistic information capture. Experiments on the controlled synergy task and real-world benchmarks demonstrate consistent improvements over existing multimodal contrastive methods, with notable gains on tasks dominated by synergistic interactions. Code is released at https://github.com/brightest66/HRIL.
CommentsAccepted at NeurIPS 2026