发表机构
Sapienza University of Rome(罗马大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于伪隔离分数的非参数离群值检测框架,实现概率定义与几何分类,并集成到K-means中形成ODK-means,通过理论与实证验证其有效性。
AI 中文摘要
离群值检测是数据处理中的一个基本挑战,对统计建模、机器学习和探索性数据分析中的稳健性具有关键影响。然而,现有的方法很少提供一种通用的、与领域无关的离群值定义,通常依赖于缺乏统计解释的启发式修剪配额。为解决这一问题,我们提出了一种基于伪隔离离群值分数的非参数框架。该分数能够实现与误报率$\alpha$相关的正式概率定义,并扩展为内部和外部离群值的严格几何分类。我们证明该机制可以无缝嵌入任何基于目标的聚类框架中,以识别特定于聚类的离群值。在此,我们将其集成到$K$-均值中,创建ODK-means。该框架的推断能力和拓扑性质通过形式化命题在理论上以及广泛的模拟和方法论教程在实证上进行了探索,突出了所提出的离群值检测逻辑的实际可操作性和直观吸引力。
英文摘要
Outlier detection is a fundamental challenge in data processing, with critical implications for robustness across statistical modeling, machine learning and exploratory data analysis. However, existing proposals rarely offer a universal, domain-agnostic definition of an outlier, often relying on heuristic trimming quotas that lack a statistical interpretation. To address this, we propose a nonparametric framework built on a pseudo-isolation outlier score. This score enables a formal, probabilistic definition of an anomaly tied to a false-alarm rate $α$, which extends into a rigorous geometric classification of internal and external outliers. We show that this mechanism seamlessly embeds into any objective-based clustering framework to identify cluster-specific outliers. Here, we integrate it into $K$-means to create ODK-means. The inferential capabilities and topological properties of this framework are explored both theoretically through formal propositions and empirically through extensive simulations and methodological tutorials, highlighting the practical actionability and intuitive appeal of the proposed outlier detection logic.