arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26974stat.MLcs.LGcs.NAmath.NAmath.STstat.TH

为何不应使用高斯核

Why not to use the Gaussian kernel

  • Lappeenranta–Lahti University of Technology LUT(拉彭兰塔-拉赫蒂理工大学(LUT大学))
  • Newcastle University(纽卡斯尔大学)

机构由 AI 辅助整理,请以论文原文为准。

Toni Karvonen, Chris J. Oates

AI总结:

本文论证高斯核因存在过度自信、数值病态等缺陷应避免使用,指出核心问题是核的解析性,进而提出应尽量避免使用解析核。

AI中文摘要:

核函数在回归、分类等任务中用于度量相似度或相关性。高斯核,又称平方指数核、径向基函数核,是高斯过程回归中最常用的核函数之一。本文认为应尽量避免使用高斯核,且绝不能将其作为默认选择,该观点基于两项证明高斯核极其脆弱的结果:其一,高斯核会产生小到不切实际的条件方差,若用该方差量化预测不确定性,几乎必然会出现灾难性的过度自信;其二,小方差与数值病态条件相伴而生,因此实际使用高斯核时需要加入 nugget 项等技巧,这些技巧会有效修改底层的回归或分类模型。这些问题由高斯核的非自然平滑性导致,并非首个注意到该事实的研究,问题的核心并非高斯形式本身,而是核函数的解析性,本文更广泛地提出应尽量避免使用解析核;对于平稳核而言,解析性本质上等价于谱密度的指数衰减。

英文摘要:

Kernels measure similarity or correlation in tasks such as regression and classification. The Gaussian kernel, other names of which include squared exponential and radial basis function kernel, is one of the most popular in Gaussian process regression. We argue that the Gaussian kernel is best avoided and should never be used as a default. The argument rests on two results demonstrating that the Gaussian kernel is extremely brittle. First, the Gaussian kernel gives rise to a conditional variance that is unrealistically small. If the variance is used to quantify predictive uncertainty, catastrophic overconfidence is almost inevitable. Second, a small variance goes hand in hand with numerical ill-conditioning, so that to use the Gaussian kernel in practice requires tricks such as nugget terms that effectively modify the underlying regression or classification model. These problems are caused by the unnatural smoothness of the Gaussian kernel, a fact we are far from the first to take notice of. The problem is not the Gaussian form itself but the analyticity of the kernel: Our argument is more broadly that analytic kernels are best avoided. For stationary kernels analyticity is essentially equivalent to an exponential decay of the spectral density.

↑