arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于分布排斥聚类的分布行列式点过程

Distributional Determinantal Point Process for Repulsive Clustering of Distributions

Khai Nguyen, Yang Ni, Elizabeth Juarez-Colunga, Peter Mueller

arXiv 2607.21847首次发表:更新:

发表机构

Department of Statistics and Data Sciences, University of Texas at Austin; Department of Mathematics, University of Texas at Austin; Department of Biostatistics and Informatics, University of Colorado Anschutz(统计与数据科学系,德克萨斯大学奥斯汀分校; 数学系,德克萨斯大学奥斯汀分校; 生物统计学与信息学系,科罗拉多大学安舒茨分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出分布行列式点过程(dDPP)用于分布排斥聚类,通过L系综与SW核构建,证明其有效性。基于此提出分布值随机划分模型,用于单细胞基因表达和人类癫痫数据推断,得到可解释且分离良好的聚类。

AI 中文摘要

我们引入分布行列式点过程(dDPP)作为一种新型排斥点过程,其原子是概率分布而非实空间中的点。dDPP通过在分布之间使用切片瓦瑟斯坦(SW)核的L系综构建。我们证明了它作为一个定义良好的点过程的有效性。在离散设置中,对于从分布原子中独立同分布采样得到的L系综、相关核及其行列式的插件估计器,我们得出了集中结果。利用该框架,我们通过排斥广义贝叶斯混合模型提出了一个分布值随机划分模型。该模型在混合测度的原子上放置dDPP先验,并基于SW距离定义广义似然。为了总结后验推断,我们开发了一种决策理论方法,在分层最优传输效用函数下作为贝叶斯规则报告混合测度的点估计。我们将所提出的框架用于单细胞基因表达数据和人类癫痫数据的推断,产生了可解释且分离良好的聚类,反映了数据中有意义的结构。

英文摘要

We introduce the distributional determinantal point process (dDPP) as a novel repulsive point process whose atoms are probability distributions rather than points in a real space. The dDPP is constructed via an L-ensemble with a sliced Wasserstein (SW) kernel between distributions. We show its validity as a well-defined point process. In the discrete setting, we derive concentration results for plug-in estimators of the L-ensemble, the correlation kernel, and their determinants given i.i.d. samples from the distributional atoms. Leveraging this framework, we propose a distribution-valued random partition model by way of a repulsive generalized Bayesian mixture model. The model places a dDPP prior over the atoms of the mixing measure and defines a generalized likelihood based on SW distance. To summarize posterior inference, we develop a decision-theoretic approach to report a point estimate of the mixing measure as a Bayes rule under a hierarchical optimal transport utility function. The latter is a natural choice given that the mixing measure is itself a distribution over distributions. We use the proposed framework for inference with single-cell gene expression data and human epilepsy data, producing interpretable and well-separated clusters that reflect meaningful structure in the data.

Comments61 pages, 15 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑