AI 中文总结
研究为何贪心搜索能产生最优聚类结果,通过分析真实与经验分布嵌入近似误差及贪心搜索与划分拟阵的关系,为‘聚类即分布’聚类目标提供近似最优保证,解释其能发现任意形状、密度和大小聚类的原因。
AI 中文摘要
许多现有聚类方法基于集合导向定义,依赖点对点相似性函数找相似点,对紧凑聚类有效,对不规则形状、密度或大小不同的聚类效果差。近期的‘聚类即分布’(CaD)聚类通过贪心搜索将每个聚类视为从未知分布生成的独立同分布点集,能发现一般类型聚类,效果优于谱聚类且无需特征分解,但缺乏理论分析。本文从两方面分析:一是分析真实与经验分布嵌入的近似误差;二是表明实现CaD聚类目标的贪心搜索可映射到划分拟阵从而产生贪心最优性。这为CaD聚类目标提供了近似最优保证,遗憾由近似误差控制。这是首次解释为何当估计的聚类嵌入忠实地近似潜在聚类分布时,通过贪心搜索的CaD聚类能发现任意形状、密度和大小的聚类(而所有基于集合导向的聚类方法都未能发现)。
英文摘要
Many existing clustering methods are designed based on a set-oriented definition---a cluster is a set of similar points---relying a point-to-point similarity function to find similar points. This works well for compact clusters, but clustering performance can deteriorate badly when cluster shapes are irregular, and densities or sizes vary between clusters. Recent `Cluster-as-Distribution' (CaD) clustering has been shown to discover these generic types of clusters in practice by treating each cluster as a set of independent and identically distributed points generated from some unknown distribution via a greedy search, achieving a clustering objective equivalent to that of Spectral Clustering, but with better clustering outcomes without eigen-decomposition. However, a theoretical analysis of this phenomenon is still lacking. Our analyses are from two angles. First, we analyze the approximation error between the true and empirical distribution embeddings. Second, we show that the greedy search employed to achieve the CaD clustering objective can be mapped to a partition matroid---yielding greedy optimality. These yield a near-optimality guarantee for the CaD clustering objective, with regret controlled by the approximation error. This is the first analysis that explains why CaD clustering via greedy search can discover clusters of arbitrary shapes, densities and sizes (where all set-oriented clustering methods have failed to discover) when the estimated cluster embeddings faithfully approximate the underlying cluster distributions.