流式细胞术高维成分数据统计方法的比较:对对数比率变换和差异丰度检验的批判性视角
Comparison of statistical methods for high-dimensional compositional data from flow cytometry: A critical perspective on log-ratio transformation and differential abundance testing
浏览论文内容
中文总结 AI 辅助
本研究比较了edgeR/TMM、CLR + limma、ANCOM-BC2等流式细胞术高维成分数据的DA分析方法,通过模拟显示ANCOM-BC2和CLR + limma更适用,CODAK/PERMANOVA可作为整体成分偏移的初筛。
中文摘要 AI 辅助
流式细胞术生成的是本质上的成分计数数据:观测到的细胞群计数被约束为等于所获取事件的总数,这排除了对绝对细胞丰度的直接推断。尽管存在这一约束,大多数流式细胞术的差异丰度(DA)分析仍依赖于最初为RNA测序开发的方法,如带有加权 trimmed mean of M 值(TMM)归一化的 edgeR,却未充分认识到所得估计值的条件性质。在此,我们针对流式细胞术DA分析中成分性的统计意义提出批判性视角。我们比较了三种针对细胞群的DA检验方法(edgeR/TMM、结合线性建模的中心对数比率变换(CLR + limma)、经偏差校正的微生物组成分分析2(ANCOM-BC2))的理论基础和解释范围,并通过涵盖7种场景、两种样本量(每组n=50和n=200)的综合模拟研究评估其性能。我们还评估了基于核的成分数据分析(CODAK)作为整体成分差异的全局综合筛选检验,其通过Aitchison距离上的PERMANOVA实现。我们的结果显示,在考虑样本量、零膨胀、离散度、细胞群间相关性以及假发现率(FDR)和灵敏度的情况下,为微生物组数据分析开发的方法ANCOM-BC2和CLR + limma比edgeR/TMM更适合流式细胞术数据的细胞群DA分析,而CODAK/PERMANOVA为整体成分偏移提供了稳健的初筛。
英文摘要
Flow cytometry generates inherently compositional count data: observed cell population counts are constrained to sum to the total number of acquired events, which precludes direct inference about absolute cellular abundance. Despite this constraint, most differential abundance (DA) analyses in cytometry rely on methods originally developed for RNA sequencing, such as edgeR with trimmed mean of M-values (TMM) normalization, without fully acknowledging the conditional nature of the resulting estimates. Here, we offer a critical perspective on the statistical implications of compositionality for flow cytometry DA analysis. We compare the theoretical foundations and interpretational scope of three per-population DA testing methods (edgeR/TMM), centered log-ratio transformation with linear modeling (CLR + limma), and analysis of composition of microbiomes with bias correction-2 (ANCOM-BC2), and evaluate their performance through a comprehensive simulation study spanning seven scenarios at two sample sizes (n = 50 and n = 200 per group). We additionally evaluate compositional data analysis using kernels (CODAK) as a global, omnibus screening test for overall compositional differences, implemented via PERMANOVA on Aitchison distances. Our results show that, considering sample size, zero-inflation, dispersion, and correlation among cell populations in addition to false discovery rate (FDR) and sensitivity, the methods developed for microbiome data analysis, ANCOM-BC2 and CLR + limma, emerge as better-suited methods for per-population DA analysis of flow cytometry data than edgeR/TMM, while CODAK/PERMANOVA offers a robust first-stage screen for overall compositional shifts.
发表机构
- National Cancer Institute/National Institutes of Health(美国国立癌症研究所/美国国立卫生研究院)
机构由 AI 辅助整理,请以论文原文为准。