arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

特征选择粒度的实证研究

An Empirical Study of Feature Selection Granularity

Muhammad Rajabinasab, Arthur Zimek

arXiv 2607.24145首次发表:更新:

发表机构

University of Southern Denmark(南丹麦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究从算法设计角度探讨特征选择,对比传统全局选择与贪婪递归消除设计,用五种算法进行实证研究,发现贪婪方法虽计算成本高,但能提高特征选择质量,验证了维度诅咒对缓解方式的影响。

AI 中文摘要

特征选择旨在为给定数据集识别最具信息性和相关性的特征。现有研究多聚焦于开发新算法、提出新评估指标和框架或对现有方法进行基准测试。本文从算法设计角度审视特征选择。传统算法全局计算特征重要性分数并一步选出顶级特征,这引发问题:信息少或有噪声的特征会掩盖更相关特征的真正重要性吗?递归策略是否更好?为此,用五种特征选择算法进行实证研究,在传统全局选择和贪婪递归消除设计下实现各算法,分析算法选择对一系列评估指标的影响。结果表明贪婪方法虽计算成本高,但几乎总能提高特征选择质量。

英文摘要

Feature selection aims to identify the most informative and relevant features for a given dataset, either in terms of capturing the underlying data structure and distribution better, or with respect to the performance on a downstream task. Existing research in this area has largely focused on developing novel algorithms (in both supervised and unsupervised settings), proposing new evaluation metrics and frameworks, or benchmarking the performance of existing methods. In this work, we examine feature selection through an algorithmic design perspective. Conventional feature selection algorithms typically compute feature importance scores globally across the entire feature set and then select the top-ranked features in a single step. However, this approach raises a critical question: Can the presence of less informative (or noisy) features mask or obscure the true importance of other, more relevant features? In other words, would a recursive strategy, where features are removed one by one while re-evaluating importance at each step, yield different and potentially better results than the standard global ranking approach? To answer this question, we conduct an extensive empirical study using five diverse feature selection algorithms. We implement each algorithm under both the conventional global selection design and the greedy recursive elimination design. We then analyze the impact of this algorithmic choice, both individually for each method and collectively across all methods, on a range of standard feature selection evaluation metrics. The empirical evaluation results show that the greedy approach improves the overall feature selection quality almost consistently, albeit on the expense of higher computational cost, supporting our initial expectation that the curse of dimensionality also obscures the ways of mitigating it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑