发表机构
Van Lang University(范朗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究单细胞聚类中深度表示学习是否必要,通过九个聚类管道在十个真实数据集上的诊断基准测试及相关分析,集成多种测试方法,揭示不同模式,确定主要方差贡献者,提供数据集感知和计算意识决策框架。
AI 中文摘要
单细胞核糖核酸测序(scRNA-seq)是精准医疗工作流程的基础技术,有助于实现联合国关于健康与福祉的可持续发展目标3,无监督聚类是将原始表达矩阵转化为可解释细胞群体的分析步骤。从业者面临一个反复出现的工程决策:额外的深度表示阶段是否值得其计算和调优成本,还是经典主成分分析(PCA)管道已经足够?我们通过对十个真实数据集(90 - 5685个细胞,19046 - 41480个基因,4 - 11种细胞类型)上的九个聚类管道进行诊断基准测试来解决这个问题,并在七个数据集上进行了部分scVI V2专门比较。该协议集成了Optuna超参数搜索、重复运行鲁棒性、Friedman/Wilcoxon-Holm/TOST测试以及Sobol全序敏感性分析。对比自动编码器实现了最高的平均调整兰德指数(0.7872),但经Holm校正的测试并未确定其优于最强基线。每个数据集的分析揭示了三种可重复的模式:概率变分自动编码器(VAE)变体在最小数据集上有帮助,深度自动编码器在具有多批次或多种类型结构的中等规模数据上获胜,而当线性投影已经捕获主要变异时,经典PCA管道仍然具有竞争力。Sobol指数确定学习率($S_T = 0.70$)和潜在维度($S_T = 0.56$)是主要的方差贡献者,表明应在何处分配有限的调优预算。因此,贡献是一个支持可持续医疗分析的生物医学人工智能管道的数据集感知和计算意识决策框架,而不是普遍优越性声明。
英文摘要
Single-cell ribonucleic acid sequencing (scRNA-seq) is a foundational technology for precision-medicine workflows that contribute to United Nations Sustainable Development Goal 3 on Good Health and Well-being, and unsupervised clustering is the analytical step that turns raw expression matrices into interpretable cell populations. Practitioners therefore face a recurring engineering decision: is an additional deep representation stage worth its compute and tuning cost, or do classical principal component analysis (PCA) pipelines already suffice? We address this question with a diagnostic benchmark of nine clustering pipelines on ten real datasets (90-5,685 cells, 19,046-41,480 genes, 4-11 cell types), augmented by a partial scVI V2 specialized comparison on seven datasets. The protocol integrates Optuna hyperparameter search, repeated-run robustness, Friedman/Wilcoxon-Holm/TOST testing, and Sobol total-order sensitivity analysis. The contrastive autoencoder achieved the highest mean Adjusted Rand Index (0.7872), but Holm-corrected tests did not establish dominance over the strongest baselines. Per-dataset analysis reveals three reproducible regimes: probabilistic variational autoencoder (VAE) variants help on the smallest datasets, deep autoencoders win on mid-scale data with multi-batch or many-type structure, and classical PCA pipelines remain competitive when linear projection already captures the dominant variation. Sobol indices identify learning rate ($S_T=0.70$) and latent dimensionality ($S_T=0.56$) as the dominant variance contributors, indicating where limited tuning budgets should be allocated. The contribution is therefore a dataset-aware and compute-conscious decision framework for biomedical AI pipelines supporting sustainable healthcare analytics, rather than a universal superiority claim.
Comments13 pages, 6 figures. Accepted at ISRSD 2026