arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28756stat.MLcs.LG

潜在空间中深度聚类的选择性推断

Selective Inference for Deep Clustering in Latent Spaces

Eina Mizui, Tomohiro Shiraishi, Shunichi Nishino, Ichiro Takeuchi

首次发表
浏览论文内容

中文总结 AI 辅助

针对深度聚类中潜在空间聚类结果的统计检验问题,提出基于固定预训练编码器的选择性推断框架,控制选择偏差并实现有效检验,合成与基因组实验验证了其功效与可靠性。

中文摘要 AI 辅助

深度聚类是一种强大的方法,通过在聚类之前学习低维潜在表示来发现高维数据中的有意义结构。尽管其在经验上取得了成功,但评估所得聚类的统计可靠性仍然具有挑战性。在同一数据上测试发现的聚类会引入选择偏差,并使经典p值失效。选择性推断(SI)为纠正这种偏差提供了一个原则性框架,但现有方法侧重于直接在观测特征上进行的聚类。在这项工作中,我们为使用固定预训练编码器的深度聚类开发了一个选择性推断框架。关键挑战在于,聚类分配是通过从原始数据空间到潜在空间的非线性变换确定的,这导致选择过程比传统聚类复杂得多。我们的方法提供了一种计算上可行的方式来解释这一过程,并能够对潜在空间中识别出的聚类之间的差异进行有效的统计检验。合成实验表明,所提出的方法控制了第一类错误率,同时比有效但保守的基线方法具有更高的功效,基因组应用表明,它可以在适当考虑选择偏差的同时识别显著的聚类差异。我们的框架为量化深度聚类发现结构的统计可靠性提供了一种原则性方法。

英文摘要

Deep clustering is a powerful approach for discovering meaningful structures in high-dimensional data by learning a low-dimensional latent representation prior to clustering. Despite its empirical success, assessing the statistical reliability of the resulting clusters remains challenging. Testing discovered clusters on the same data induces selection bias and invalidates classical $p$-values. Selective inference (SI) provides a principled framework for correcting this bias, but existing methods focus on clustering performed directly on the observed features. In this work, we develop an SI framework for deep clustering with a fixed pretrained encoder. The key challenge is that cluster assignments are determined through a nonlinear transformation from the original data space to the latent space, resulting in a substantially more complex selection process than in conventional clustering. Our method provides a computationally tractable way to account for this process and enables valid statistical testing of differences between clusters identified in the latent space. Synthetic experiments demonstrate that the proposed method controls the Type I error rate while achieving higher power than valid but conservative baselines, and genomic applications show that it can identify significant cluster differences while appropriately accounting for selection bias. Our framework provides a principled approach to quantifying the statistical reliability of structures discovered by deep clustering.

发表机构

  • Nagoya University(名古屋大学)
  • RIKEN(理化学研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑