最大编码率降低用于分布外泛化的局限性
On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation
浏览论文内容
中文总结 AI 辅助
本文揭示最大编码率降低(MCR²)在分布外泛化中的两大局限:其目标可导致完全预测失败,且结合不变性原理仍无法消除,需新假设或原则确保跨环境稳定预测。
中文摘要 AI 辅助
大量工作致力于使深度学习目标、表示和架构具有可解释性,旨在提高学习系统在多样化实际应用中的安全性、鲁棒性和泛化能力。最近提出的最大编码率降低($\mathrm{MCR}^{2}$)为学习类别子流形的结构化、判别性表示提供了一个有前景的信息论框架,并启发了可解释的白盒架构。然而,我们观察到$\mathrm{MCR}^{2}$在分布偏移下可能完全失效,这促使我们研究其分布外(OOD)泛化的局限性。我们确立了$\mathrm{MCR}^{2}$在OOD泛化方面的两个局限。首先,仅凭$\mathrm{MCR}^{2}$目标本身就可能允许完全预测失败:完全基于不稳定环境特征的表示可以达到全局编码最优,但在相关性反转后完全失败,尽管存在一个完全稳定的特征。这个精确最优的例子包括训练中不可能出现的测试输入。即使每个可能的测试输入在训练中也可能出现,编码质量可以任意接近最优,而预测误差可以任意接近100%。其次,直接纳入广泛成功的风险不变最小化(IRM)和风险外推(REx)所依据的不变性原理并不能消除这种失败。失败的表示在训练环境之间具有相同的最优编码算子,这表明共享的编码最优性并不能确保稳定的预测。因此,$\mathrm{MCR}^{2}$的可靠OOD保证需要额外的新假设或学习原则,以建立跨环境的稳定预测关系。
英文摘要
Substantial efforts have been devoted to making deep learning objectives, representations, and architectures interpretable, with the goal of improving the safety, robustness, and generalisation of learning systems in diverse real-world applications. The recently proposed maximal coding rate reduction ($\mathrm{MCR}^{2}$) offers a promising information-theoretic framework for learning structured, discriminative representations of class-wise submanifolds and has inspired interpretable white-box architectures. However, we observe that $\mathrm{MCR}^{2}$ can completely fail under distribution shift, motivating our study of its out-of-distribution (OOD) generalisation limits. We establish two limitations of $\mathrm{MCR}^{2}$ for OOD generalisation. First, the $\mathrm{MCR}^{2}$ objective alone can admit complete prediction failure: a representation based entirely on unstable environmental features can achieve the global coding optimum yet fail completely after correlation reversal, despite an available perfectly stable feature. This exact-optimum example includes test inputs that cannot occur during training. Even when every possible test input can also occur during training, coding quality can be arbitrarily close to optimal while prediction error is arbitrarily close to 100%. Second, directly incorporating the invariance principle underlying widely successful invariant risk minimisation (IRM) and risk extrapolation (REx) does not eliminate this failure. The failing representation admits the same optimal coding operator across training environments, showing that shared coding optimality does not ensure stable prediction. Reliable OOD guarantees for $\mathrm{MCR}^{2}$ therefore require additional new assumptions or learning principles that establish stable predictive relationships across environments.
发表机构
- University of Sheffield(谢菲尔德大学)
机构由 AI 辅助整理,请以论文原文为准。