arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于非参数分布筛选的聚类降维

Cluster-Based Dimensionality Reduction by Nonparametric Distributional Screening

Sanoja Jha, Rishikesh Muralimohan, Praveen Athauda Arachchi, Abhishek Bhattacharjee

arXiv 2609.06854首次发表:更新:

AI 中文总结

针对带聚类划分的高维数据,提出基于多样本Kolmogorov-Smirnov统计量的非参数筛选方法,保留可解释的判别性坐标,并建立理论保证与一致性准则。

AI 中文摘要

我们考虑针对高维观测数据的降维问题,这些观测数据附带一个给定的划分,将其分为两个或多个聚类。目标不是构造低秩投影,而是保留原始坐标中可解释的子集,该子集能够保留区分各聚类的分布信息。对于每个坐标,所提出的方法通过多样本Kolmogorov-Smirnov分离统计量来比较聚类特定的经验分布函数。我们将由此得到的边际聚类支持形式化,并建立了所有坐标上的同时有限样本集中性、假包含和遗漏的显式界,以及当最小分布分离度主导高维随机误差时的精确支持恢复。我们还量化了由于使用未调整的检验水平所导致的维度膨胀,并给出了一个控制族系错误率的版本。在条件充分性条件下,确定筛选保留了全数据后验聚类概率、互信息和贝叶斯风险;另一个结果刻画了对不完美估计聚类标签的稳健性。该方法对严格递增的坐标变换具有不变性,并且可以保留主成分可能丢弃的低方差聚类信号。我们进一步开发了平均对偶信息,这是一个结合变换后划分一致性与聚类相关坐标结构覆盖的准则,并推导了其基本性质和一致性。模拟实验说明了理论、所选坐标的可解释性,以及聚类导向筛选与方差导向投影之间的区别。

英文摘要

We consider dimensionality reduction for high-dimensional observations accompanied by a supplied partition into two or more clusters. The objective is not to construct a low-rank projection, but to retain an interpretable subset of the original coordinates that preserves the distributional information distinguishing the clusters. For each coordinate, the proposed procedure compares the cluster-specific empirical distribution functions through a several-sample Kolmogorov-Smirnov separation statistic. We formalize the resulting marginal cluster support and establish simultaneous finite-sample concentration over all coordinates, explicit bounds for false inclusions and omissions, and exact support recovery when the minimum distributional separation dominates the high-dimensional stochastic error. We also quantify the dimension inflation induced by using an unadjusted testing level and give a familywise-error-controlled version. Under a conditional sufficiency condition, sure screening preserves the full-data posterior cluster probabilities, mutual information, and Bayes risk; an additional result characterizes robustness to imperfectly estimated cluster labels. The procedure is invariant to strictly increasing coordinate transformations and can retain low-variance cluster signals that principal components may discard. We further develop average dual information, a criterion combining partition agreement after transformation with structural coverage of cluster-relevant coordinates, and derive its basic properties and consistency. Simulations illustrate the theory, the interpretability of the selected coordinates, and the distinction between cluster-directed screening and variance-directed projection.

Comments32 Pages, 3 tables, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑