arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于表示学习的子模信息测度目标:方差与分离视角的理解

Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective

Rishabh Iyer, Truong Pham, Anay Majee

arXiv 2607.27660首次发表:更新:

发表机构

The University of Texas at Dallas; Adobe(德克萨斯大学达拉斯分校; 奥多比公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建统一理论框架揭示子模信息测度(SIMs)的几何与统计特性,通过控制合成实验验证,为选择设计基于SIMs的表示学习目标提供原则性指导。

AI 中文摘要

子模信息测度(SIMs)近年已成为表示学习与多模态学习的强大框架,其中SCORE框架证明SIMs可作为监督对比学习的有效目标。尽管其经验上取得成功,但不同子模信息测度诱导的几何与统计特性仍鲜为人知。本研究构建了将SIMs与表示学习、统计模式识别经典概念关联的统一理论框架,证明总信息(TI)目标表征类内结构:图割TI恢复类内方差,对数行列式TI恢复广义方差与协方差体积,设施选址TI诱导感知不平衡的分离,强调稀有类与易混淆类;互信息(MI)目标捕捉互补的类间结构概念:图割MI与质心分离、费希尔式判别密切相关,对数行列式MI通过马氏距离捕捉协方差感知的分离,设施选址MI测量近模态表征重叠。本研究通过控制合成实验验证这些理论表征,实验独立改变方差、协方差、类别不平衡、类别分离及多模态重叠,所有设置下的经验行为均与提出的理论高度匹配。本研究结果首次提供了子模信息测度的统一几何与统计理解,为选择和设计基于SIMs的表示学习目标提供了原则性指导。

英文摘要

Submodular Information Measures (SIMs) have recently emerged as a powerful framework for representation learning and multimodal learning. In particular, the SCORE framework~\cite{majee2024score} demonstrated that SIMs can serve as effective objectives for supervised contrastive learning. Despite their empirical success, however, the geometric and statistical properties induced by different submodular information measures remain poorly understood. In this work, we develop a unified theoretical framework connecting SIMs to classical concepts in representation learning and statistical pattern recognition. We show that Total Information (TI) objectives characterize intra-class structure: Graph Cut TI recovers within-class variance, LogDet TI recovers generalized variance and covariance volume, and Facility Location TI induces imbalance-aware separation that emphasizes rare and confusable classes. We further show that Mutual Information (MI) objectives capture complementary notions of inter-class structure: Graph Cut MI is closely related to centroid separation and Fisher-style discrimination, LogDet MI captures covariance-aware separation through Mahalanobis distance, and Facility Location MI measures nearest-mode representational overlap. We validate these theoretical characterizations using controlled synthetic experiments that independently vary variance, covariance, class imbalance, class separation, and multimodal overlap. Across all settings, the empirical behavior closely matches the proposed theory. Our results provide the first unified geometric and statistical understanding of submodular information measures and offer principled guidance for selecting and designing SIM-based objectives for representation learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑