arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高维聚类的非参数假设检验及其在单细胞RNA数据中的应用

Nonparametric Hypothesis Testing of High-dimensional Clustering With Application to Single-cell RNA Data

Yifan Dai, Di Wu, Yufeng Liu

arXiv 2609.05683首次发表:更新:

发表机构

University of North Carolina; University of Michigan(北卡罗来纳大学; 密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对高维聚类显著性检验中高斯假设不可靠的问题,提出基于对数凹分布的非参数方法SigClust-LCP,结合分数匹配估计器,在控制第一类错误和保持功效的同时,避免单细胞数据中的虚假亚聚类。

AI 中文摘要

单细胞RNA测序研究通常使用聚类来定义假定的细胞类型和细胞状态,然而观察到的分离可能源于采样变异性,而非真正的生物异质性。本文研究了高维数据中此类聚类结构的正式显著性检验。现有的SigClust方法通过在高斯单聚类零假设下进行蒙特卡洛模拟来评估聚类显著性,但这一假设对于归一化基因表达数据和其他非高斯设置可能不可靠。我们提出了SigClust-LCP,一种非参数扩展方法,使用对数凹分布对单个聚类进行建模。为了使该方法在中高维度下计算可行,我们受近期生成建模思想的启发,开发了一种用于对数凹投影的分数匹配估计器。我们为该估计器及其在聚类显著性检验中的使用建立了理论保证。模拟表明,在一系列单峰和混合分布中,SigClust-LCP比现有方法更可靠地控制第一类错误,同时保持有竞争力的功效。在对水螅细胞的单细胞RNA测序分析中,该方法避免了注释细胞群体内的虚假亚聚类,并支持跨谱系和体轴区域的生物学上有意义的分离。

英文摘要

Single-cell RNA sequencing studies routinely use clustering to define putative cell types and cell states, yet the observed separation may arise from sampling variability rather than genuine biological heterogeneity. This paper studies formal significance testing of such clustering structure in high-dimensional data. Existing SigClust methods assess clustering significance through Monte Carlo simulation under a Gaussian single-cluster null, but this assumption can be unreliable for normalized gene expression data and other non-Gaussian settings. We propose SigClust-LCP, a nonparametric extension that models a single cluster by a log-concave distribution. To make this approach computationally feasible in moderate to high dimensions, we develop a score-matching estimator for log-concave projection inspired by recent generative modeling ideas. We establish theoretical guarantees for the estimator and for its use in clustering significance testing. Simulations show that SigClust-LCP controls Type-I error more reliably than existing methods across a range of unimodal and mixture distributions while retaining competitive power. In a single-cell RNA sequencing analysis of Hydra cells, the method avoids spurious subclusters within annotated cell populations and supports biologically meaningful separation across lineages and body-axis regions.

Comments29 pages, 4 figures and 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑