arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23117cs.LGcs.AI

白化反转层级结构:白化嵌入的范数度量了什么

Whitening Inverts the Hierarchy: What the Norm of a Whitened Embedding Measures

Mohammed Ahnouch, Lotfi Elaachak

首次发表
浏览论文内容

中文总结 AI 辅助

本文揭示白化嵌入平方范数并非高斯似然,而是语义非典型性的马氏度量,解释了其有效性及校准失败原因。

中文摘要 AI 辅助

将基础模型嵌入进行白化,并将其平方范数用作免训练似然代理,这一做法的动机是观察到白化后的坐标通常近似标准正态分布。我们证明,这一观察源于投影中心极限定理,因此并不蕴含高斯联合分布。在多个编码器和三种训练目标下,我们发现白化半径相对于高斯参考存在系统性过度离散,即使与具有相同均值和协方差的分布克隆相比也是如此。我们进一步表明,经验范数统计量与理论范数统计量之间通常报告的一致性,是样本内白化的代数结果,并不构成高斯性的证据。我们识别出这一行为背后的机制:白化反转了编码器的谱层级,将平方范数的贡献转向主要编码噪声的近退化方向。在这些方向上,主导变异性由单个依赖于输入的尺度控制。我们根据两个矩估计该尺度,并利用它无需额外自由参数即可预测不相交谱半部分之间的交叉依赖性。这些结果表明,白化平方范数更适合解释为语义非典型性的马氏度量,而非对数似然。这一解释既说明了其实际有效性,也解释了其校准失败:该统计量能够与非参数密度估计以及在不同目标下训练的编码器一致地排序和检测非典型样本,而高斯尾部阈值可能不准确达数个数量级。

英文摘要

Whitening a foundation-model embedding and using its squared norm as a training-free likelihood surrogate is motivated by the observation that whitened coordinates often appear approximately standard normal. We show that this observation follows from the projection central limit theorem and therefore does not imply a Gaussian joint distribution. Across multiple encoders and three training objectives, we find systematic over-dispersion of the whitened radius relative to the Gaussian reference, including against distributional clones with identical mean and covariance. We further show that the commonly reported agreement between empirical and theoretical norm statistics is an algebraic consequence of in-sample whitening and does not constitute evidence for Gaussianity. We identify the mechanism behind this behavior: whitening reverses the encoder's spectral hierarchy, shifting the contribution to the squared norm toward near-degenerate directions that encode predominantly noise. In these directions, the dominant variability is governed by a single input-dependent scale. We estimate this scale from two moments and use it to predict, without additional free parameters, the cross-dependence between disjoint spectral halves. These results indicate that the squared whitened norm is better interpreted as a Mahalanobis measure of semantic atypicality than as a log-likelihood. This interpretation explains both its practical effectiveness and its calibration failures: the statistic can rank and detect atypical samples consistently with nonparametric density estimates and across encoders trained with different objectives, while Gaussian tail thresholds can be inaccurate by orders of magnitude. etc.

发表机构

  • Université Paris 1(巴黎第一大学)
  • Faculty of Science and Technology of Tangier(丹吉尔科学与技术学院)
  • Abdelmalek Essaadi University(阿卜杜勒马利克·埃萨迪大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑