AI 中文总结
提出自动深度局部中心聚类(A-DLCC),利用β-集成局部深度识别局部中心,结合自适应合并准则,无需参数调整即可自动确定簇数并产生可解释聚类结果。
AI 中文摘要
聚类是一种无监督学习技术,将未标记的数据划分为若干组。现有的大多数方法需要用户指定参数,如簇的数量或邻域大小。相反,我们提出了自动深度局部中心聚类(A-DLCC),一种完全数据驱动的方法,消除了数值参数调整。A-DLCC使用$\eta$-集成局部深度来识别稳定的范例点,即那些在多个局部性层级上始终处于中心的点,称为局部中心,并根据其代表性进行排序。每个局部中心诱导一组相似的点,组间相似性通过所提出的非参数度量——组级局部相似性来测量。为了指导合并,我们引入了图论中的瓶颈路径思想,这构成了我们自适应合并准则的基础。基于该准则,我们设计了一个单一的凝聚规则,其中一组要么被其比自身更能到达的邻居吸收,要么与双方都认为比自身背景更可到达的邻居结合,每次合并还额外要求由比配置模型零假设所预期的更强的接触来承载。该规则自动估计簇的数量并决定何时停止合并。在合成数据和真实数据上的实验表明,A-DLCC无需参数调整即可产生可解释的聚类结果。
英文摘要
Clustering is an unsupervised learning technique that partitions unlabeled data into groups. Most existing methods require user-specified parameters, such as the number of clusters or neighborhood size. Conversely, we propose automatic depth-based local center clustering (A-DLCC), a fully data-driven method that eliminates numerical parameter tuning. A-DLCC uses the $β$-integrated local depth to identify stable exemplars, points consistently central across multiple locality levels, termed local centers, which are ranked by their representativeness. Each local center induces a group of similar points, with group-level similarity measured by a proposed nonparametric metric called group-level local similarity. To guide merging, we incorporate the bottleneck path idea from graph theory, which forms the basis of our adaptive merging criterion. Based on this criterion, we design a single agglomeration rule in which a group is either absorbed by a neighbor it reaches better than itself or bonded to a neighbor that both sides find more reachable than their own background, every merge being additionally required to be carried by a contact stronger than a configuration-model null expects. The rule automatically estimates the number of clusters and decides when to stop merging. Experiments on synthetic and real data show that A-DLCC produces interpretable clustering results without parameter tuning.