全局协方差池化中矩阵对数归一化的正交多项式逼近
Orthogonal Polynomial Approximation for Matrix Log Normalization in Global Covariance Pooling
浏览论文内容
中文总结 AI 辅助
该研究针对全局协方差池化中矩阵对数归一化的数值不稳定问题,提出用正交多项式逼近实现无分解的矩阵对数归一化,在多数据集上验证了其速度与精度优势。
中文摘要 AI 辅助
全局协方差池化(GCP)通过捕获二阶特征统计量来改进深度网络,尤其对细粒度识别有效。由于协方差矩阵位于对称正定(SPD)流形上,在欧几里得分类器前需进行归一化步骤,最忠实的选择是矩阵对数(MLN-COV),它将SPD流形映射到其切空间;但实际应用中因基于特征分解的梯度数值不稳定,被矩阵平方根替代。本文表明,该不稳定性是光谱计算对数的产物,而非对数本身的问题。用协方差矩阵的有限多项式逼近对数可消除两次计算中的特征分解:所有操作均成为通用矩阵乘法(GEMM),梯度在归一化前协方差的光谱支撑上保持有界,且不会出现不稳定的1/(λ_i-λ_j)项。关键要素是均值特征值预归一化,其将光谱中心移至1附近,远离对数的奇点,辅以标量后补偿以闭式返回log(A)的奇异部分。本文推荐的归一器是8阶切比雪夫展开,通过三项矩阵递推计算,反向传播则用匹配的逆递推;同时研究了勒让德、拉盖尔、泰勒和帕德展开作为对照,以分离基函数和目标函数的作用。在三个细粒度基准和ImageNet-1k上,无分解对数相比光谱对数及所替代的平方根近似,速度更快且精度更高,在匹配基函数和阶数下,对数目标优于平方根目标,证实增益来自忠实的黎曼映射而非更优的多项式族。
英文摘要
Global Covariance Pooling (GCP) improves deep networks by capturing second-order feature statistics, and is especially effective for fine-grained recognition. Because covariance matrices live on the Symmetric Positive Definite (SPD) manifold, a normalization step is required before the Euclidean classifier. The faithful choice is the matrix logarithm (MLN-COV), which maps the SPD manifold to its tangent space; in practice it was abandoned in favour of the matrix square root because its eigendecomposition-based gradient is numerically unstable. We show that this instability is an artifact of computing the logarithm spectrally, not of the logarithm itself. Approximating the logarithm with finite polynomials in the covariance matrix removes the eigendecomposition from both passes: every operation becomes a General Matrix Multiplication (GEMM), the gradient stays bounded on the spectral support of the pre-normalized covariance, and the unstable 1/(lambda_i-lambda_j) term never appears. The key ingredient is a mean-eigenvalue pre-normalization that centres the spectrum near 1, away from the singularity of log, with a scalar post-compensation that returns the singular part of log(A) in closed form. Our recommended normalizer is a degree-8 Chebyshev expansion evaluated by a three-term matrix recurrence, with a matching reverse recurrence for the backward pass; Legendre, Laguerre, Taylor and Pade expansions are studied as controls that isolate the roles of the basis and of the target function. On three fine-grained benchmarks and ImageNet-1k the decomposition-free logarithm is both faster and more accurate than the spectral logarithm and than the square-root approximations it replaces, and at matched basis and degree the log target beats the square-root target, confirming that the gain comes from the faithful Riemannian map rather than from a better polynomial family.
发表机构
- NeuronAITree(神经元AI树)
- Cambridge University(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。