arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高维线性判别分析中的双重下降现象

Double Descent in High-dimensional Linear Discriminant Analysis

Yonghe Lu, Han Lin Shang, Yanrong Yang, Kehan Zhao

arXiv 2609.19061首次发表:更新:

发表机构

The Australian National University; Macquarie University(澳大利亚国立大学; 麦考瑞大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文利用随机矩阵理论推导了高维线性判别分析(LDA)在欠参数化和过参数化情形下的渐近误分类风险,完整刻画了其双重下降风险曲线,并通过模拟和ARCENE数据集实验验证了理论结果。

AI 中文摘要

现代机器学习模型中观察到的双重下降现象挑战了经典的偏差-方差权衡,表明当模型复杂度超过插值阈值时,预测误差可以再次下降。尽管这一现象已在深度神经网络和其他高容量模型中得到广泛研究,但其对经典统计方法的影响仍未被充分理解。本文研究了线性判别分析(LDA)中的双重下降现象,LDA是一种用于二分类的基本分类方法。利用随机矩阵理论,我们在维度(p)和样本量(n)满足(p/n \to \u03b3)的高维情形下推导了LDA的渐近误分类风险。具体而言,当(\u03b3 \u2208 (0,1))时,我们获得了具有一般总体协方差矩阵的LDA风险的显式渐近表达式;当(\u03b3 \u2208 (1,\u221e))时,我们获得了在各向同性协方差模型下基于Moore-Penrose伪逆的LDA分类器的显式渐近表达式。这些结果共同提供了LDA在欠参数化和过参数化两种情形下双重下降风险曲线的完整刻画。模拟研究和在ARCENE癌症分类数据集上的实验与理论预测高度一致。

英文摘要

The double-descent phenomenon observed in modern machine learning models has challenged the classical bias-variance trade-off by showing that prediction error can decrease again as model complexity exceeds the interpolation threshold. Although this phenomenon has been extensively studied for deep neural networks and other high-capacity models, its implications for classical statistical methods remain less well understood. In this paper, we investigate double descent in linear discriminant analysis (LDA), a fundamental classification method for binary classification. Using random matrix theory, we derive the asymptotic misclassification risk of LDA in the high-dimensional regime where the dimension (p) and sample size (n) satisfy (p/n \to γ). Specifically, we obtain an explicit asymptotic expression for the LDA risk with a general population covariance matrix when (γ\in (0,1)), and for the Moore-Penrose pseudo-inverse LDA classifier under an isotropic covariance model when (γ\in (1,\infty)). Together, these results provide a complete characterization of the double-descent risk curve of LDA across both the under- and over-parameterized regimes. Simulation studies and experiments on the ARCENE cancer classification dataset demonstrate close agreement with the theoretical predictions.

Comments28 pages, 17 figures, 1 table

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑