arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

模型无关的特征选择:基于LOCO引导的自适应小批量采样

Model-Agnostic Feature Selection via LOCO-Guided Adaptive Minipatch Sampling

Xuhui Liu, Lili Zheng

arXiv 2609.24126首次发表:更新:

发表机构

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LAMPS,一种模型无关的集成特征选择框架,利用LOCO得分自适应采样,在高维下实现精确选择,优于现有方法。

AI 中文摘要

黑盒机器学习模型日益能够提供强大的预测能力,但从这些模型中提取有用信息(例如一组重要特征)仍然具有挑战性。现有的模型无关方法主要估计特征重要性或对其推断,而非直接选择特征,而许多特征选择方法是模型特定的或依赖于模型X假设。我们提出了LOCO引导的自适应小批量采样(LAMPS),这是一种模型无关的集成框架,使用任何黑盒回归算法作为其基学习器来选择对预测响应重要的特征。基学习器仅需产生预测,无需自身执行特征选择。LAMPS在小批量集成框架内运作,该框架对观测值和特征进行子采样,使得留一协变量(LOCO)特征重要性得分易于计算。它自适应地将小批量采样集中在具有高LOCO得分的特征上,同时保持探索性。经过几次迭代后,所得的采样概率迅速将信号特征与噪声特征区分开,从而能够通过简单的阈值化进行选择。我们证明,在基预测模型平均训练充分的情况下,LAMPS在高维设置中实现了精确的特征选择。在合成数据和真实数据上的大量实验表明,LAMPS优于最先进的特征选择方法,尤其在存在相关特征时表现尤为突出。

英文摘要

Black-box machine learning models increasingly deliver strong predictions, but extracting useful information from them, such as a set of important features, remains challenging. Existing model-agnostic methods primarily estimate feature importance or conduct inference on it rather than directly selecting features, whereas many feature selection methods are model-specific or rely on the model-X assumption. We introduce LOCO-guided Adaptive Minipatch Sampling (LAMPS), a model-agnostic ensemble framework that uses any black-box regression algorithm as its base learner to select features important for predicting the response. The base learner need only produce predictions and need not perform feature selection itself. LAMPS operates within a minipatch ensemble framework that subsamples both observations and features, allowing leave-one-covariate-out (LOCO) feature importance scores to be easily computed. It adaptively concentrates minipatch sampling on features with high LOCO scores while maintaining exploration. The resulting sampling probabilities rapidly separate signal from noise features after a few iterations, enabling selection through simple thresholding. We establish that LAMPS achieves exact feature selection in high-dimensional settings, provided that the base predictive models are sufficiently well trained on average. Extensive experiments on synthetic and real data show that LAMPS outperforms state-of-the-art feature selection methods, with particularly strong performance in the presence of correlated features.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑