多重共线性无关的非欧几里得响应特征筛选:一种因子调整方法
Multicollinearity-agnostic feature screening for non-Euclidean responses: a factor adjusted approach
浏览论文内容
中文总结 AI 辅助
针对高维非欧几里得响应中多重共线性导致边际筛选失效的问题,提出因子调整的Fréchet确信独立筛选方法,通过公共因子与特异成分分离,保证筛选性质并验证有效性。
中文摘要 AI 辅助
在高维设置中,多重共线性是一个普遍存在的问题,它会显著损害基于边际Fréchet回归的特征筛选方法的性能。当超高维预测变量遭受多重共线性时,非欧几里得响应的特征筛选变得不可靠,因为特征特定的信号可能被共享的潜在因子掩盖。为了减轻这种影响,我们提出了一种因子调整的Fréchet确信独立筛选程序。该方法首先从预测变量中恢复潜在公共因子,然后通过每个特征的特异成分在公共因子之外贡献的增量Fréchet决定系数来评估每个特征。在正则条件下,我们建立了可行筛选效用的均匀逼近率,并证明了确信筛选和确信排序性质。广泛的数值实验为我们的方法的有效性和效率提供了令人信服的经验支持,特别是在具有高度相关协变量的场景中。我们进一步通过两个具有代表性的非欧几里得数据集(ADNI数据集和死亡率数据集,两者均具有分布值响应)说明了我们方法的实际性能。
英文摘要
In high-dimensional settings, multicollinearity is a pervasive issue that can substantially impair the performance of feature screening methods based on marginal Fréchet regression. Feature screening for non-Euclidean responses becomes unreliable when ultrahigh-dimensional predictors suffer from multicollinearity, because feature-specific signals may be masked by shared latent factors. To mitigate this effect, we propose a Factor adjusted Fréchet sure independence screening procedure. The method first recovers latent common factors from the predictors and then evaluates each feature by the incremental Fréchet coefficient of determination contributed by its idiosyncratic component beyond the common factors. Under regularity conditions, we establish uniform approximation rates for the feasible screening utilities and prove the sure screening and sure ranking properties. Extensive numerical experiments provide compelling empirical support for the validity and effectiveness of our approach, particularly in scenarios with highly correlated covariates. We further illustrate the practical performance of our method through two representative non-Euclidean datasets: the ADNI dataset and the mortality dataset, both with distribution-valued responses.
发表机构
- Tsinghua University(清华大学)
- Beijing Normal University(北京师范大学)
- Shanghai University of Finance and Economics(上海财经大学)
机构由 AI 辅助整理,请以论文原文为准。