发表机构
Northeast Normal University; School of Mathematics and Statistics(东北师范大学; 数学与统计学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于广义斯坦引理的充分降维框架,通过构建交叉矩矩阵与奇异值分解恢复中心子空间,兼具无需线性条件、可利用未标记数据等优势,在标记稀缺、高噪声场景中性能优于现有方法。
AI 中文摘要
充分降维(SDR)旨在寻找预测变量的最小子空间,该子空间能够捕获响应变量的完整条件分布,这一子空间被称为中心子空间(CS)。当响应变量为多元时,该问题会变得极具挑战性,尤其是在样本量有限的情况下。现有方法存在不同局限性:逆回归方法依赖强分布假设和矩阵求逆,其多响应扩展方法存在严重的切片稀疏问题;前向回归方法依赖计算密集型迭代平滑,其计算成本随响应维度增加而增长;基于深度学习的方法需要大量标记数据。为规避这些缺陷,本文提出一种基于广义斯坦引理的SDR框架,该方法构建多元响应变量与预测变量边缘得分函数之间的交叉矩矩阵,并通过其奇异值分解恢复CS。所提方法不依赖线性条件,避免了矩阵求逆和迭代平滑,且在有未标记数据时可加以利用。本文在标准正则条件下为所提估计量建立了收敛保证,还提出一种实用的秩选择算法以估计CS的维度。大量模拟研究和实际数据应用表明,所提方法在多种场景下均优于现有方法,尤其在中等维度、标记数据稀缺且噪声水平高的场景中表现突出。
英文摘要
Sufficient dimension reduction (SDR) seeks the minimal subspace of the predictors that captures the full conditional distribution of the response, which is known as the central subspace (CS). When the response is multivariate, the problem becomes considerably more challenging, particularly when the sample size is limited. Existing methods face different limitations:inverse regression approaches rely on strong distributional assumptions and matrix inversion, and their multi-response extensions suffer from severe slice sparsity; forward regression methods depend on computationally intensive iterative smoothing whose cost grows with the response dimension; and deep learning-based approaches demand large amounts of labeled data. To circumvent these shortcomings, we propose an SDR framework based on the generalized Stein's lemma. Our method constructs a cross-moment matrix between the multivariate response and the marginal score function of the predictors, and recovers the CS via its singular value decomposition. The proposed method does not rely on the linearity condition, avoids matrix inversion as well as iterative smoothing, and can leverage unlabeled data when available. We establish convergence guarantees for the proposed estimator under standard regularity conditions. Moreover, we propose a practical rank-selection algorithm to estimate the dimension of the CS. Extensive simulation studies and a real data application demonstrate that the proposed methods consistently outperform existing approaches across a variety of settings, particularly in moderate-dimensional, label-scarce scenarios with high noise levels.