AI 中文总结
提出基于密度的留一法影响分数用于无监督离群点检测,用LBFP估计器定义分数,有闭式更新,计算高效。研究污染下分数,模拟显示其性能及计算时间,信用卡欺诈应用表明该方法在大型真实数据集上效果良好。
AI 中文摘要
我们提出了一种基于密度的留一法影响分数用于无监督离群点检测。动机是离群点自然与概率密度非常小的区域相关,但直接的留一法密度重新拟合计算量可能过大。我们使用线性混合频率多边形(LBFP)估计器并定义一个分数,该分数比较在一个观测值处的全样本拟合密度与移除该观测值后得到的拟合密度,同时保持网格和带宽固定。所得统计量测量观测值自身位置处的相对密度扰动。对于LBFP估计器,该分数有精确的闭式更新,所以无需为每个观测值重新拟合密度估计器。这在保持直接密度解释的同时使该方法对大样本计算高效。我们研究了污染情况下的分数,表明常规正密度观测值和污染驱动观测值有不同的渐近阶。在广泛的污染模型上的模拟说明了这些理论情况,相对于标准基准显示出有竞争力的性能,并记录了计算时间。一个有29个变量的信用卡欺诈应用表明该方法在大型真实数据集上效果良好。
英文摘要
We propose a density-based leave-one-out influence score for unsupervised outlier detection. The motivation is that outliers are naturally associated with regions of very small probability density, but direct leave-one-out density refitting can be computationally prohibitive. We use the Linear-Blend Frequency Polygon (LBFP) estimator and define a score that compares the full-sample fitted density at an observation with the fitted density obtained after removing that observation, while keeping the grid and bandwidth fixed. The resulting statistic measures a relative density perturbation at the observation's own location. For the LBFP estimator, this score has an exact closed-form update, so the density estimator does not need to be refitted for each observation. This preserves a direct density interpretation while making the method computationally efficient for large samples. We study the score under contamination and show that regular positive-density observations and contamination-driven observations have distinct asymptotic orders. Simulations over a broad range of contamination models illustrate these theoretical regimes, show competitive performance relative to standard benchmarks, and document computing time. A credit-card fraud application with 29 variables illustrates that the method works well on a large real data set.
Comments26 pages, 3 figures, supplementary material included