发表机构
Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对树模型变量重要性偏向连续变量的问题,提出添加噪声的简单去偏方法,并结合稳定性选择实现混合数据变量选择与错误发现控制。
AI 中文摘要
基于树的方法(如随机森林)产生的变量重要性分数偏向于连续预测变量而非分类预测变量。我们对此偏差进行了理论分析,并提出了一种简单的补救措施:向每个分类预测变量添加少量噪声。该修正方法在多种模拟和真实世界数据集上得到了验证,并与集成路径稳定性选择相结合,以在混合数据中执行具有错误发现控制的变量选择。
英文摘要
Variable importance scores from tree-based methods such as random forests favor continuous predictors over categorical ones. We present a theoretical analysis of this bias and propose a simple remedy: add a small amount of noise to each categorical predictor. The correction is demonstrated on a variety of simulated and real-world datasets and combined with integrated path stability selection to perform variable selection with false discovery control for mixed data.