arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01175cs.LG

重新思考信息瓶颈:标签诱导划分下的结构化分解

Rethinking the Information Bottleneck: Structured Decomposition under Label-Induced Partitions

Jingyao Zhang, Yuxuan Li, Lu Han, Ali Anaissi, Nguyen H. Tran

首次发表
浏览论文内容

中文总结 AI 辅助

针对标准信息瓶颈无法独立调节标签相关与组内信息的问题,提出基于标签诱导划分的双瓶颈公式,通过条件KL分解控制干扰变异,在低数据分类和密集预测中取得一致改进。

中文摘要 AI 辅助

标准信息瓶颈(IB)正则化通过单一标量 I(Z;X) 约束表示,隐式地将所有信息视为同质。然而,单一的全局压缩控制将标签相关结构与残余的组内变异耦合在一起,而非独立调节它们的分配,导致干扰信息得以在学习的表示中存留。例如,在医学影像应用中,残余变异通常源于采集条件、背景因素或受试者特有的外观。这一问题在数据有限的环境中尤为突出,模型倾向于过拟合此类变异,从而阻碍泛化。尽管现有正则化方法能够稳定训练、控制容量或塑造表示几何,但它们并未明确地将干扰样变异与任务支持结构分离。为解决这一局限,我们基于标签诱导划分从结构化视角重新审视信息瓶颈,其中条件级结构与组内信息扮演不同角色。这引出了一个双瓶颈公式:一个标准 KL 项控制全局信息容量,而一个条件 KL 项针对组内信息。我们证明条件 KL 可精确分解为组内信息项和先验失配项,解释了其与设计的一致性。通过使用单纯形结构的条件先验,该方法提供了可控的潜在几何,并能无缝集成到现有流程中。在分类和分割上的实验表明,在低数据分类中增益最为显著,并在密集预测基准上实现了一致的改进。

英文摘要

Standard information bottleneck (IB) regularization constrains representations via a single scalar I(Z;X), implicitlytreating all information as homogeneous. However, a single global compression control couples label-relevant structurewith residual within-condition variation, rather than regulating their allocation independently, allowing nuisanceinformation to persist in learned representations. For example, in medical imaging applications, residual variation oftenstems from acquisition conditions, background factors, or subject-specific appearance. This issue becomes particularlypronounced in data-limited settings, where models tend to overfit such variation, hindering generalization. While existingregularization methods can stabilize training, control capacity, or shape representation geometry, they do not explicitlyseparate nuisance-like variation from task-supporting structure. To address this limitation, we revisit IB from a structuredperspective based on a label-induced partition, where condition-level structure and within-condition information playdistinct roles. This leads to a dual-bottleneck formulation: a standard KL term controls global information capacity, while aconditional KL term targets within-condition information. We show that the conditional KL admits an exact decompositioninto a within-condition information term and a prior-mismatch term, explaining its alignment with the design objective.With a simplex-structured conditional prior, the method provides controllable latent geometry and integrates seamlesslyinto existing pipelines. Experiments on classification and segmentation show the clearest gains in low-data classificationand consistent improvements across dense prediction benchmarks.

发表机构

  • The University of Sydney(悉尼大学)

机构由 AI 辅助整理,请以论文原文为准。

↑