arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多元随机森林在特征选择中的理论性质及其在面部形态-基因检测中的应用

Theoretical Properties of Multivariate Random Forest in Feature Selection and its Application to Facial Morphology-Gene Detection

Yangsheng Wang, Samruddhi Thakar, Anton Schick, Guifang Fu

arXiv 2607.21880首次发表:更新:

AI 中文总结

研究为多变量结果联合特征选择建理论基础,将MRF的PVIM作为高维特征选择工具,证明其一致性,通过面部形态GWAS展示其实用性,模拟表明其性能优于其他方法,还提出新模拟框架。

AI 中文摘要

本文为多变量结果的联合特征选择建立了理论基础,将多元随机森林(MRF)基于排列的变量重要性度量(PVIM)定位为高维特征选择的原则性工具。我们建立了MRF的首个一致性保证,表明在温和正则条件下,随着样本量趋于无穷,它以趋于1的概率保留所有真正有影响的特征。采用不完全U统计量纳入三层随机性。PVIM是一种联合筛选方法,能处理多重共线性等问题。通过人类面部形态的全基因组关联研究证明了MRF的实用性,它识别出多个新基因座和相互作用中心。广泛模拟表明MRF能准确识别有影响的信号,优于其他方法。此外,还提出了一个新的模拟框架。

英文摘要

This work establishes a theoretical foundation for joint feature selection with multivariate outcomes, positioning the permutation-based variable importance measure (PVIM) of multivariate random forests (MRF) as a principled tool for high-dimensional feature selection. We establish the first consistency guaranty for MRF, showing that it retains all truly influential features with probability tending to one as the sample size grows to infinity under mild regularity conditions. Incomplete U-statistics is employed to incorporate three layers of randomness: subsampling of subjects for training each tree, subsampling of features at each split, and permutation of each feature for the out-of-bag (OOB) samples. Unlike independence-based screening that evaluates each feature in isolation, PVIM is a joint screening approach that accounts for multicollinearity, nonlinear, high-order interactions, and subject heterogeneity via ensemble aggregation. Moreover, we demonstrate the practical utility of MRF through a genome-wide association study (GWAS) of human facial morphology (with 2,342 subjects and 453,273 SNPs), where MRF identifies several novel loci and interaction hubs that extend prior findings. Extensive simulations show that MRF accurately identifies truly influential signals while producing parsimonious feature sets with well controlled false selection rates, outperforming canonical correlation analysis (CCA) and several other independence multivariate screening approaches. In addition, we also propose a novel simulation framework, including image outcomes, that more closely mimic the intricate nature of real-world data and provide rigorous testbeds for machine learning research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑