AI 中文总结
该研究针对幂律各向异性下的核岭回归,推导高维 regime 中核谱与泛化误差的渐近尖锐表达式,明确各向异性对学习曲线、偏差与方差转变及泛化性能的影响。
AI 中文摘要
我们针对各向异性高斯数据下的核岭回归展开研究,其中针对多项式内积核,输入协方差以指数α≥0的幂律形式衰减。我们在多项式高维 regime n=Θ(d^κ)中推导了核谱与泛化误差的渐近尖锐表达式,揭示了各向异性如何重塑学习曲线。对于弱各向异性(0<α<1),问题仍为有效高维,保留了各向同性情形的部分特征,同时在其他方面偏离:方差仍在整数样本复杂度κ∈ℕ处达到峰值,但这些峰值随α增大逐渐衰减;同时,对于与数据主方向高度对齐的目标,偏差在分数样本复杂度处下降,使偏差转变与插值峰值解耦。对于强各向异性(α>1),问题的有效维度为常数,方差不再依赖样本量,在无岭插值下趋于平稳,或在固定岭罚则下以显式速率消失。偏差经历由目标衰减率决定的尖锐转变:低于阈值时,学习是突变而非渐进;高于阈值时,偏差以幂律衰减,恢复经典源与容量速率。我们最终将这些结果应用于单索引目标,展示索引与数据主方向的对齐如何决定各向异性对学习的影响。综上,我们的结果阐明了输入几何如何塑造核特征并从根本上影响其泛化性能。
英文摘要
We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent $α\geq 0$ for polynomial inner-product kernels. We derive asymptotically sharp expressions for the kernel spectrum and the generalization error in the polynomial high-dimensional regime $n=Θ(d^κ)$, revealing how anisotropy reshapes the learning curves. For weak anisotropy ($0<α<1$), the problem remains effectively high-dimensional and retains some features of the isotropic case, while departing from it in others: the variance still peaks at integer sample complexities $κ\in\mathbb{N}$, but these peaks are progressively damped as $α$ grows; meanwhile, for targets strongly aligned with the data's principal directions, the bias drops at fractional sample complexities, decoupling the bias transitions from the interpolation peaks. For strong anisotropy ($α> 1$), the effective dimension of the problem is constant, and the variance stops depending on sample size altogether, plateauing under ridgeless interpolation or vanishing at an explicit rate under fixed ridge penalty. The bias undergoes a sharp transition governed by the target's decay rate: below a threshold, learning is abrupt rather than gradual; above it, the bias decays as a power law that recovers the classical source and capacity rates. We finally specialize these results to single-index targets, showing how the alignment of the index with the data's principal directions determines the effect of anisotropy on learning. Together, our results clarify how the input geometry shapes the kernel features and fundamentally impacts its generalization properties.