发表机构
Auburn University; Augusta University(奥本大学; 奥古斯塔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对稀疏支持向量机,提出基于线性规划与对偶变量的去偏估计方法,实现高维下特征推断与FDR控制变量选择,并通过模拟和实际数据验证。
AI 中文摘要
利用副本对称的高维刻画,我们在样本量与特征数成比例增长的情况下,为稀疏支持向量机开发了一个推断框架。主要挑战在于非光滑的铰链损失,这阻碍了为光滑分类损失设计的去偏论证的直接应用。我们通过将$L_1$惩罚支持向量机(SVM)表示为线性规划,并利用其对偶变量识别铰链损失的次梯度,从而克服了这一困难。这产生了一个计算上可行的去偏估计量,在比例渐近机制下,其坐标渐近服从高斯分布。由此得到的分布特征为单个特征提供了置信区间和假设检验,并实现了错误发现率控制的变量选择。广泛的模拟研究考察了在各种协方差结构(包括强相关设计)下的校准、功效和变量选择性能。对高维乳腺癌基因表达数据的分析说明了所提出的推断如何将统计显著的特征与原始稀疏SVM选择的变量区分开来。
英文摘要
Using a replica-symmetric high-dimensional characterization, we develop an inferential framework for sparse support vector machines when the sample size and number of features grow proportionally. The main challenge is the nonsmooth hinge loss, which prevents direct application of debiasing arguments developed for smooth classification losses. We overcome this difficulty by representing the $L_1$-penalized support vector machine (SVM) as a linear program and identifying the hinge-loss subgradient through its dual variables. This yields a computationally accessible debiased estimator whose coordinates are asymptotically Gaussian under the proportional asymptotic regime. The resulting distributional characterization provides confidence intervals and hypothesis tests for individual features and enables false-discovery-rate-controlled variable selection. Extensive simulations examine calibration, power, and variable-selection performance under a range of covariance structures, including strongly correlated designs. An analysis of high-dimensional breast cancer gene-expression data illustrates how the proposed inference can distinguish statistically significant features from variables selected by the original sparse SVM.
Comments7 figures