AI 中文总结
针对高维数据的异聚类结构问题,提出一种频率框架下的局部谱聚类方法,可同时识别特征组并估计对应样本聚类结构,经模拟和应用验证了其优越性。
AI 中文摘要
经典聚类方法通常假设所有有信息的特征都支持观测值的单个潜在划分,该假设对于现代高维数据可能过于严格——不同的特征子集可能编码不同的相似性概念并诱导异质样本划分,而部分特征可能不包含有意义的聚类信息。我们开发了一种用于局部聚类的频率框架,该框架可同时识别特征组并估计与每组相关的样本聚类结构。我们的方法通过标签不变的聚类矩阵表示每个样本划分,并根据特征共享的聚类结构对特征进行分组,从而将局部聚类重新表述为特征分组(即聚类的聚类)问题。在异质子高斯混合模型下,我们构建了特定于特征的高斯核相似矩阵,并提出了一种基于聚类矩阵优化准则的局部谱聚类程序。所提方法避免了显式似然指定和贝叶斯后验计算,可适应异质特征分布,且允许无信息特征存在。大量模拟和应用进一步证明了该方法的实用性和优越性。
英文摘要
Classical clustering methods typically assume that all informative features support a single latent partition of the observations. This assumption can be overly restrictive for modern high-dimensional data, where different subsets of features may encode distinct notions of similarity and induce heterogeneous sample partitions, while some features may contain no meaningful clustering information. We develop a frequentist framework for local clustering that simultaneously identifies feature groups and estimates the sample clustering structure associated with each group. Our approach represents each sample partition by a label-invariant clustering matrix and groups features according to their shared clustering structures, thereby reformulating local clustering as a feature-grouping, or clustering-of-clusterings, problem. Under a heterogeneous sub-Gaussian mixture model, we construct feature-specific Gaussian-kernel similarity matrices and propose a local spectral clustering procedure based on a clustering-matrix optimization criterion. The proposed method avoids explicit likelihood specification and Bayesian posterior computation, accommodates heterogeneous feature distributions, and permits the presence of non-informative features. Extensive simulations and applications further demonstrate the practical utility and superiority of the proposed approach.