arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14421cs.ITmath.IT

依赖性、压缩与协同:多模态学习的统一信息论视角

Dependency, Compression, and Synergy: A Unified Information-Theoretic View of Multimodal Learning

发表机构西南财经大学计算机与人工智能学院 · 北京人形机器人创新中心(X-Humanoid) · 中国科学院深圳先进技术研究院
另 1 家 · 查看机构详情
  • School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics(西南财经大学计算机与人工智能学院)
  • Beijing Humanoid Robot Innovation Center (X-Humanoid)(北京人形机器人创新中心(X-Humanoid))
  • Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
  • University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Liangjian Wen, Linjie Li, Jiang Duan, Yong Dai, Jianzhuang Liu, Zhao Kang

首次发表
浏览论文内容

中文总结 AI 辅助

本综述提出统一信息论框架,通过MI、IB和PID连接多模态学习原理,回顾170项研究,提出广义多模态信息拉格朗日量坐标系,并指出开放挑战。

中文摘要 AI 辅助

近年来,多模态基础模型的进展加剧了理解不同模态如何共享、保留和补充信息的需求。互信息(MI)、信息瓶颈(IB)和部分信息分解(PID)提供了互补的视角,然而现有研究往往将它们视为孤立的工具。本综述提出了一种信息论视角,将这些原理连接为多模态信息处理逐步精炼的观点:MI刻画模态间依赖性,IB解释压缩下面向任务的信息保留,PID将保留的信息分解为冗余性、独特性和协同性。我们回顾了170项近期研究(2018年至2026年)和12项基础工作,围绕四个挑战组织多模态学习:跨模态对齐、信息高效融合、交互类型刻画以及向多模态基础模型的扩展。我们不以应用领域作为主要分类轴,而是将医疗健康、机器人技术、推荐系统、情感计算和无线通信解释为这些信息原理的经验验证。超越分类学,我们在一个统一的信息论坐标系——广义多模态信息拉格朗日量——中组织现有多模态范式,它们占据精确或近似的参数角落,其未占据区域命名了文献尚未构建的候选方法族。我们进一步讨论了新兴多模态基础模型如何大规模实例化这些原理,并识别了开放挑战,包括高维设置中的可扩展信息估计、跨信息论方法的标准化评估、多模态PID的组合复杂性,以及从事后信息分析向信息感知的多模态学习的转变。

英文摘要

Recent advances in multimodal foundation models have intensified the need to understand how different modalities share, preserve, and complement information. Mutual Information (MI), the Information Bottleneck (IB), and Partial Information Decomposition (PID) provide complementary perspectives, yet existing studies often treat them as isolated tools. This survey presents an information-theoretic perspective connecting these principles as progressively refined views of multimodal information processing: MI characterizes inter-modal dependency, IB explains task-oriented information preservation under compression, and PID decomposes preserved information into redundancy, uniqueness, and synergy. We review 170 recent studies (2018--2026) and 12 foundational works, organizing multimodal learning around four challenges: cross-modal alignment, information-efficient fusion, interaction-type characterization, and scaling to multimodal foundation models. Rather than using application domains as primary taxonomy axes, we interpret healthcare, robotics, recommendation systems, affective computing, and wireless communications as empirical validations of these information principles. Beyond taxonomy, we organize existing multimodal paradigms within a single information-theoretic coordinate system -- the Generalized Multimodal Information Lagrangian -- in which they occupy exact or approximate parameter corners, and whose unoccupied regions name candidate method families the literature has not yet built. We further discuss how emerging multimodal foundation models instantiate these principles at scale and identify open challenges including scalable information estimation in high-dimensional settings, standardized evaluation across information-theoretic methods, combinatorial complexity of multimodal PID, and the transition from post-hoc information analysis toward information-aware multimodal learning.

↑