arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

泊松子空间聚类:聚焦计数数据的核心要素

Poisson Subspace Clustering: Focusing on the Essentials in Count Data

Collin Leiber, Kai Puolamäki, Heikki Mannila

arXiv 2608.23287首次发表:更新:

发表机构

Aalto University; University of Helsinki(阿尔托大学; 赫尔辛基大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对通用算法无法适配计数数据分布特性的问题,提出基于泊松等统计模型的3CPO聚类算法,可识别相关列子集提升可解释性,经多领域实验验证其聚类质量与鲁棒性。

AI 中文摘要

计数数据以非负整数值矩阵(如列联表)形式呈现,广泛存在于各类领域。对这类数据集进行聚类时,需采用特定方法,因为通用算法往往未考虑其独特的分布特性,会导致输出结果不可靠。一种有效策略是使用泊松、负二项分布等成熟统计模型。我们提出3CPO算法,这是一种基于计数数据统计可靠建模的聚类算法,除了聚类标签外,它还能识别相关列的子集,提升结果的可解释性。我们提出一种简单的迭代算法,通过最大化后验概率来寻找优质聚类解,并讨论其特性。大量实验表明,该算法能在基因表达、文本、经济学等不同领域的相关子空间中定义高质量聚类。研究表明,3CPO是一种统计可靠且可解释的计数数据聚类的鲁棒解决方案,代码可在指定URL获取。

英文摘要

Count data represented as a matrix of non-negative integer values, such as contingency tables, are prevalent across diverse domains. When clustering such data sets, specific methods are required, as generic algorithms often fail to consider their unique distributional properties, leading to unreliable outputs. An effective strategy is to use well-established statistical models such as the Poisson and negative binomial distributions. We present 3CPO, a clustering algorithm based on statistically solid modeling of count data. In addition to the cluster labels, it identifies a subset of relevant columns, enhancing the interpretability of the results. We propose a simple iterative algorithm that maximizes the posterior probability to find good clustering solutions and discuss its properties. Extensive experiments demonstrate its ability to define high-quality clusters within associated subspaces for various data domains, ranging from gene expressions and texts to economics. Our findings suggest that 3CPO is a robust solution for clustering count data in a statistically sound and interpretable manner. Our code is available at https://github.com/collinleiber/3CPO.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑