无模型且分布鲁棒的特征筛选与高维异质数据中的错误发现控制
Model-free and Distributionally Robust Feature Screening with False Discovery Control for High-Dimensional Heterogeneous Data
浏览论文内容
中文总结 AI 辅助
本文提出基于Copula散度的无模型特征筛选方法CD-Screen及错误发现控制程序CD-FDR,适用于高维异质数据,理论保证筛选一致性和FDR控制,模拟与实证均优于传统方法。
中文摘要 AI 辅助
本文提出了一种针对高维异质数据集的模型无关特征筛选框架,该框架基于一种新颖的分布鲁棒依赖性度量——Copula散度。所提出的筛选方法名为CD-Screen,解决了现有特征筛选方法的关键局限性,如限制性建模假设和对异质特征分布的敏感性。CD-Screen根据特征的Copula散度对特征进行排序,而不依赖于特定的回归模型或分布假设。此外,我们引入了CD-FDR,一种数据驱动的程序来控制错误发现,确保准确高效的特征选择。理论分析确立了CD-Screen的确定筛选和秩一致性性质,以及CD-FDR对错误发现率的渐近控制。广泛的模拟研究表明,与传统筛选方法相比,我们的方法在各种场景下均表现出优越性能。此外,一项关于美国股票回报与通货膨胀关系的真实数据分析展示了我们方法的实际应用,并提供了关于各行业对经济变化反应的描述性证据。
英文摘要
In this paper, we propose a model-free feature screening framework tailored for high-dimensional and heterogeneous datasets, based on a novel distributionally robust dependence measure termed Copula Divergence. The proposed screening method, named CD-Screen, addresses critical limitations of existing feature screening methods, such as restrictive modeling assumptions and sensitivity to heterogeneous feature distributions. CD-Screen ranks features according to their Copula Divergence without relying on a specific regression model or distributional assumptions. Additionally, we introduce CD-FDR, a data-driven procedure to control false discoveries, ensuring accurate and efficient feature selection. Theoretical analyses establish the sure screening and rank consistency properties of CD-Screen, along with asymptotic control of the false discovery rate by CD-FDR. Extensive simulation studies demonstrate the superior performance of our methods compared to traditional screening approaches across diverse scenarios. Furthermore, a real data analysis of the relationship between stock returns and inflation in the United States {illustrates the practical use of our method and provides descriptive evidence on} sector-specific responses to economic changes.
发表机构
- University of Georgia(佐治亚大学)
- Pennsylvania State University(宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。