GCTAg:面向生物库规模农业队列的可扩展混合模型分析
GCTAg: scalable mixed-model analysis for biobank-scale agricultural cohorts
浏览论文内容
中文总结 AI 辅助
本研究通过优化GCTA的内存和CPU瓶颈,并引入降秩Woodbury矩阵方法,实现了生物库规模农业队列中高效且精确的混合线性模型关联分析。
中文摘要 AI 辅助
全基因组关联研究旨在识别与性状相关的基因组变异。使用全基因组关系矩阵的混合线性模型关联(MLMA)方法,如GCTA中实现的方法,功能强大但计算成本高昂。在此,我们消除了GCTA中的关键内存和CPU瓶颈,将REML内存使用量减少近75%,并在生物库规模队列中将MLMA加速数个数量级,同时保持精确性。我们进一步利用映射队列中的亲缘关系,通过降秩Woodbury矩阵方法,在控制基因组膨胀的同时,带来进一步的数个数量级性能提升。原生即时显性重编码还消除了中间文件上的缓慢I/O操作,使得在大型农业队列中能够进行高效的加性和显性MLMA分析。
英文摘要
Genome-wide association studies identify genomic variants associated with traits. Mixed linear model association (MLMA) methods using a whole-genome relationship matrix, such as those implemented in GCTA, are powerful but computationally expensive. Here, we remove key memory and CPU bottlenecks in GCTA, reducing REML memory usage by nearly 75% and substantially accelerating MLMA by orders of magnitude in biobank-scale cohorts while preserving exactness. We further exploit relatedness in the mapping cohort through a reduced-rank Woodbury matrix approach, delivering further orders of magnitude performance gains with controlled genomic inflation. Native on-the-fly dominance recoding also eliminates slow I/O-operations on intermediate files, enabling efficient additive and dominance MLMA analyses in large agricultural cohorts.
发表机构
- ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。